BLESS

system

t--:--:--

grid--

status

t+0.000headlineperf-max vs native, c76: whiterun +10.97%, flight +4.37%. 5 of 5 pairs each
t+0.001repeatc80 reproduces it: +12.09% and +4.04%. 1.52x and 1.22x stock dxvk
t+0.002modssame mods on both sides: the lead holds, and grows under community shaders
t+0.003strictvanilla's exact image: level in whiterun, about 5.5% behind in the flight
t+0.004latencylower on presentmon's clock. the two clocks differ. camera test next

scope

written after 5 days and 18 hours of testing, 57.7k+ loc and 2,370 benchmark runs.

I forked dxvk, the layer that translates direct3d 11 into vulkan, and tuned that fork narrowly for skyrim special edition :3 the vulkan fork had to beat skyrim in its native dx11. the bar was ultra settings at 1080p.

the final numbers come from campaign c76, the hardened build:

scenenative dx11our fork (perf-max)paired difference (95% ci)1% lowsdisplay latency
whiterun218.1 fps242.2 fps+10.97% (+10.05 to +11.89), 5 of 5 pairs170 vs 1644.6 ms vs 9.8 ms
flying traversal239.8 fps250.7 fps+4.37% (+3.70 to +5.03), 5 of 5 pairs216 vs 2044.7 ms vs 12.4 ms

average frame rate goes to us in both scenes, and 1% lows led in c76. the independent repeat (c80) put the lows tied with native, while the frame-rate gain returned at +12.09% in whiterun and +4.04% in the traversal.

on presentmon’s display-latency clock the fork reads under half of native’s figure, a number with an asterisk the present-path section explains. that same repeat had stock dxvk 3.1.1, the version the project started from, 26% behind native in whiterun and 15% behind in the traversal. measured side by side, the fork therefore runs at 1.52 times stock in whiterun and 1.22 times in the traversal.

[fig. 01]native dx11 in whiterun: 218.1 fps, 1% low 164.2, display latency 9.8 ms. ultra 1080p, rtx 3060 ti and ryzen 5 5600, median of 5 paired runs (campaign c76)
native dx11 in whiterun: 218.1 fps, 1% low 164.2, display latency 9.8 ms. ultra 1080p, rtx 3060 ti and ryzen 5 5600, median of 5 paired runs (campaign c76)
[fig. 02]our fork (perf-max) in whiterun: 242.2 fps, 1% low 170.2, display latency 4.6 ms. +11.0% fps over native, 95% ci above zero, 5 of 5 runs faster
our fork (perf-max) in whiterun: 242.2 fps, 1% low 170.2, display latency 4.6 ms. +11.0% fps over native, 95% ci above zero, 5 of 5 runs faster
[fig. 03]native dx11 flying the plains: 239.8 fps, 1% low 204.2, display latency 12.4 ms. ultra 1080p, median of 5 paired runs (campaign c76)
native dx11 flying the plains: 239.8 fps, 1% low 204.2, display latency 12.4 ms. ultra 1080p, median of 5 paired runs (campaign c76)
[fig. 04]our fork (perf-max) flying the plains: 250.7 fps, 1% low 216.2, display latency 4.7 ms. +4.4% fps over native, 95% ci above zero, 5 of 5 runs faster
our fork (perf-max) flying the plains: 250.7 fps, 1% low 216.2, display latency 4.7 ms. +4.4% fps over native, 95% ci above zero, 5 of 5 runs faster

the project is two repositories:

  • blessed-skyrim holds the benchmark harness (bench/) and the skse plugin that patches the engine (skse/skybench). every campaign plan lives there too. the tuning report holds each number and its run id, plus the long writeup this report is set from.
  • blessed-dxvk is our fork of dxvk 3.1.1 on the blessed branch, home of the translation-layer work. the constant-buffer ring and mirror, threaded front end, present path, the shader replacements and the temporal levers all live in it.

the fork delivers most of the speed. the plugin brings the engine-side patches a translation layer cannot reach, like the cascade walk and the half-rate far cascade.

the rig

it ran on one mid-range machine:

  • rtx 3060 ti, driver 596.36
  • ryzen 5 5600 (6 cores), 32gb of ddr4 ram
  • a 75 Hz monitor
  • windows 11 stripped
  • skyrim se 1.7.104 with skse 2.3.1

two scenes are the scope. whiterun parks the camera in the middle of the city, surrounded by npcs, fires and buildings. it’s cpu-bound, limited by the game’s main thread. traversal is a no-clip flight across the plains outside whiterun, gpu-bound, limited by the graphics card. two attack angles.

two profiles, this actually matters:

  • strict adds no temporal reuse. its design contrast is vanilla’s exact image, where each change is either bit-identical or checked against vanilla by a named component test.
  • perf-max allows temporal reuse. slow-moving effects (volumetric light, far shadow cascade, the water reflection cube) update every second or third frame while the camera holds steady. ⛔️ the image is not pixel-identical to vanilla.
  • bit 8 must be set in nvidia inspector, more on that later.

under strict, whiterun is level with native and the traversal trails by about 5.5%, for reasons explained near the end. under perf-max the fork takes both scenes. whenever this report says “beat native”, perf-max is meant.

day zero

performance was not the original main goal. ray tracing was.

the opening project was performant real ray tracing in skyrim se: ray-traced sun and light shadows, at least two bounces of global illumination, and reflections. the template was blessed minecraft, a minecraft mod built earlier. it runs two-bounce path tracing through voxels in about 3 to 4 ms on an rtx 3060 ti at 1080p.

nvidia rtx remix looked like the obvious tool, but we dropped it on the first evening. we decided we build our own tracer in a dxvk fork. under dxvk, skyrim’s gpu buffers are already vulkan buffers. every vertex buffer and index buffer the game creates lives inside dxvk’s memory as a VkBuffer. a ray tracer inside the fork can therefore build its acceleration structures straight from the game’s own geometry // export ❌, second copy of the world ❌ second renderer ❌. remix would have to pull every mesh out and upload it again while we get to skip it.

skyrim’s rasterizer keeps drawing everything you see directly: materials, sky, water, ui. the tracer takes over only the secondary rays, meaning shadows, bounced light and reflections.

the first night went fast:

2026-09-22 20:50bootskyrim boots on our own dxvk fork. msvc built it in 89 seconds, no patches
2026-09-22 20:55rqray queries live on skyrim's vulkan device
2026-09-22 21:28raysfirst hardware rays inside skyrim. two hits, two misses, as the geometry says. main menu at 61.1 fps
2026-09-22 23:10sunsun shadow-mask pass found (pixel shader b070feb5, pass 111 of 185) and forced to zero. we own the sun
2026-09-23 00:24lightfirst light. traced sun shadows from 600 blas and an 884-instance tlas, one ray per pixel, 0.79 ms gpu

forcing that mask to zero from inside the fork at 23:10 removed all direct sunlight from whiterun. by 00:24 the traced shadows were coming from 600 bottom-level acceleration structures and an 884-instance top level, rebuilt every frame from the game’s own buffers.

[fig. 05]whiterun untouched, with the game's own sun shadow mask in place
whiterun untouched, with the game's own sun shadow mask in place
[fig. 06]the sun's shadow mask forced to zero from inside the fork: whiterun with no direct sunlight
the sun's shadow mask forced to zero from inside the fork: whiterun with no direct sunlight

the sun direction we read from shader constant ps b2 c0 made for dramatic shadows. they also moved on their own, because that constant was not the sun. it held the shadow filter’s per-frame noise rotation, spinning around the vertical axis. the real sun sat 18 registers later at ps b2 c18, steady at (0.764, 0.085, 0.639). against the real sun the traced mask matched vanilla’s to 3.4/255 mean difference.

[fig. 07]first light: vanilla skyrim se, sun shadows from shadow maps
first light: vanilla skyrim se, sun shadows from shadow maps
[fig. 08]first light, with the wrong sun: sun shadows ray traced in our dxvk fork. the shadows followed the filter's noise rotation
first light, with the wrong sun: sun shadows ray traced in our dxvk fork. the shadows followed the filter's noise rotation
[fig. 09]the traced sun visibility behind that frame: 1 ray per pixel, 884 instances, 0.79 ms
the traced sun visibility behind that frame: 1 ray per pixel, 884 instances, 0.79 ms
[fig. 10]first light, fixed: the same vanilla frame, sun shadows from shadow maps
first light, fixed: the same vanilla frame, sun shadows from shadow maps
[fig. 11]ray-traced sun shadows in our dxvk fork from the real sun direction: stable, the game's own sun
ray-traced sun shadows in our dxvk fork from the real sun direction: stable, the game's own sun
[fig. 12]the traced sun visibility behind that frame: 1 ray per pixel, 884 instances, 0.74 ms
the traced sun visibility behind that frame: 1 ray per pixel, 884 instances, 0.74 ms
[fig. 13]skyrim's own sun shadow mask, from 2,559 shadow-map draws
skyrim's own sun shadow mask, from 2,559 shadow-map draws
[fig. 14]our ray-traced mask: 884 instances plus 70 skinned actors, 0.82 ms, no shadow maps
our ray-traced mask: 884 instances plus 70 skinned actors, 0.82 ms, no shadow maps

campaign c2 brought something great with it. inside our fork the plugin skipped the 2,559 raster shadow-cascade draws that the traced shadows replace. gpu busy time barely moved, 6.04 ms against 6.07 ms with vanilla’s shadow maps in a scene capped at 60 fps(😔). ray-traced sun shadows came in at about zero cost.

[fig. 15]global illumination v0 in whiterun, before: ray-traced sun shadows with vanilla ambient
global illumination v0 in whiterun, before: ray-traced sun shadows with vanilla ambient
[fig. 16]with ray-traced gi v0: a probe grid written into skyrim's own ambient constant buffer (0.13 ms)
with ray-traced gi v0: a probe grid written into skyrim's own ambient constant buffer (0.13 ms)
[fig. 17]the proof: the same injection with a constant red ambient
the proof: the same injection with a constant red ambient

the render-thread crisis

c3 assembled the whole stack, and frame rate fell from 61.0 to 57.7 fps, falling under skyrim’s own 60 fps cap. the gpu sat busy only 6.61 ms per frame. something on the cpu was stalling.

a custom in-driver probe found the stall. DrawIndexed was eating 4.1 ms per frame on the game’s thread. the gi readback ring lived in write-combined memory, which the gpu writes quickly but the cpu reads at a crawl. moving the ring to cached memory brought the frame back to 60.9 fps.

removing the 60 fps cap (c4) exposed a steep cliff to climb:

backendfps, uncapped
native dx11222.7
stock dxvk163.5
our fork, features off85.4
our fork, full ray-traced stack65.8

with every feature switched off, our fork ran at half the speed of stock dxvk, which is already well below native, yet the probe claimed the game thread spent only 0.20 ms inside dxvk per frame. that figure was wrong, there was no way; the probe was lying: it rounded every call duration down to whole microseconds, so a 50 ns dxvk call read as zero. the probe itself also added 4.6 ms per frame, which the 60 fps cap had been hiding. four campaigns of “dxvk’s front end is cheap” were built on a broken instrument.

a fixed probe (raw tsc ticks, calibrated, its own cost subtracted) told the truth. skyrim spends 2.1 ms per frame inside dxvk, not 0.29 ms, and our hooks were adding 2.6 ms of their own.

one morning of surgery followed:

  • the first rounds copied bones into a per-frame arena, lifting 101.1 to 109.6 fps.
  • next round named the cost. reading 3,840 bytes of skinning bones from a mapped constant buffer ran 7.6 µs per draw, about 500 MB/s, because those bytes were cold.
  • stopped reading them on the cpu. the skinning shader reads each draw’s own constant buffer by device address, reaching 118.7 fps.
  • one more round measured neutral, and we reverted it.
  • moving the gi ambient patch onto dxvk’s worker thread reached 131.0 fps, and moving the static scene capture there as well reached 143.5 fps.

the lesson written down that morning held for the whole project. what cost on skyrim’s thread was cold-cache latency, not instructions. making hooks cheaper for two rounds recovered 17 fps, while moving the same work to another thread recovered about 25 more. campaign c10 closed the morning at 139.1 fps for the full ray-traced stack, up from 101.1 at dawn. native ran 221.5 and stock dxvk 164.1. still a ways to go, but we are making progress.

numbers i can trust

after two major issues i built a real benchmark harness, and by the end of this project had produced 2,370 runs across 80 campaigns. it captured 12,640,563 frames (about 19.3 hours of gameplay) and 30 GB of data. i also keep runs that failed, which i store as more data.

bench/skybench.py launches skyrim through skse and stages the backend, whether native, stock dxvk, or a fork build, along with ini files and mods. it loads a save through the console and waits for the scene to settle. then it captures 30 seconds with presentmon and nvidia-smi and restores the game folder from a journal afterward. every file gets recorded before the harness touches it. if a run crashes, --restore replays the journal. if a file changed in a way it did not expect, the harness refuses to delete it and asks an operator.

bench/campaign.py runs paired, interleaved campaigns. each repeat walks through every configuration, native included. the next repeat walks back in reverse order. that way slow drift (a warming gpu, a background task, the room) hits all sides about equally. the first repeat serves as warm-up and after gets dropped.

[fig. 18]a paired, interleaved campaign. a, b, c are configurations
 repeat 1   a → b → c      warm-up. dropped
 repeat 2   c → b → a      reverse order
 repeat 3   a → b → c      forward again
 repeat 4   c → b → a
   …        alternating to the end of the plan

 native is one of a, b, c in every campaign
 slow drift lands on every side about equally
 headline: the average % gap to native over
           the paired runs, with a 95% range

headline results are the mean of the paired percentage differences, with a student-t 95% confidence interval. other numbers here (medians, single diagnostic runs, synthetic tests) carry a label saying what they are. one audit caught that we had used the wrong t-value for five pairs, 2.57 instead of 2.78. the right one was used from then on.

the gates grew out of failures, one at a time:

  • device loss. a run whose dxvk log shows VK_ERROR_DEVICE_LOST fails, even when presentmon kept some frames. one run lost the gpu, kept 2.3 seconds of frames, and reported a fake 228.7 fps.
  • coverage. capturing less than 85% of the window fails the run.
  • stalls. any frame interval over one second fails the run.
  • contamination. a run fails when a compiler, or any program started from a seat’s worktree or our scratch folder, ran during the capture. one campaign had nine build seats compiling in the background, and native swung from 212 to 100 fps.
  • screenshots. every run’s screenshot is classified for sun shadows, a gate born from our best near-miss. an allocator experiment reported a +18% win, but the screenshot looked like a different camera angle with an unlit wall. same angle, in fact, with the ray-traced shadows silently missing, because the allocator’s null allocation left the gpu less work to do. that number nearly snuck into the reports.

runs taken solo after an idle machine are void as well. one such run put our fork at +51% over native. and it was not noticed immediately that native’s gpu had not clocked up yet, reading 1,703 MHz at 83 W against 1,890 MHz at 132 W once settled.

the cpu war

dxvk has a traditional weakness on cpu-bound dx11 games, and skyrim is a hard case even among those. its renderer makes about 13,000 Map(DISCARD) calls per frame, 12,990 in our probe count. tens of thousands of other d3d11 calls join them, all on one thread.

there is no separate render thread; community writeups describe skyrim’s “render thread” as a second thread beside the game logic, but our thread census found exactly one os thread doing everything. Main::Update (ai, papyrus, animation, havok) and Main::Render (culling, draw submission) run one after another on it, taking about 4.4 ms per frame in whiterun. a sampled cpu profile attributed about three-quarters of that (75%) to the engine’s own Main::Render. the six job workers, meanwhile, burn two full cores spinning with no real work queued.

three weapons won the war.

the constant-buffer ring. previously each Map(DISCARD) allocated a slice from dxvk’s general allocator and emitted its own command. the ring instead hands each frame one big, persistently mapped buffer. a map becomes a pointer bump, and the “rename” rides along on data dxvk was sending anyway. we had estimated 0.5 to 0.8 ms of savings but measured 1.7 ms per frame, for +33% fps. most of the gain came from contention removed around the calls, not from the calls themselves.

the cascade walk. even when the new traced shadows replaced the raster sun cascades, the engine still walked every shadow caster on the cpu. our first fix skipped the whole walk and earned +25.6% fps. it also removed every sun shadow in the game, because the engine sets the sun’s shadow-receiver bit on geometry during that same walk. the correct fix skips only the sun’s caster queueing while keeping the walk, and keeps +20.7% with the shadows intact.

custom threaded front end. nicknamed “the cursed one”, it borrows an idea from mesa’s gallium drivers and wine’s wined3d. every eligible d3d11 call gets recorded into a lock-free ring on skyrim’s thread, at an estimated 5 to 8 ns per call. a second thread then replays it into dxvk’s real state tracking and vulkan recording. buffer writes still return at once on the game thread because allocation is already thread-safe. the design required deferred release and no atomics on the hot path. each buffer also keeps two copies of its current pointer, one for the app and one for the replay thread.

[fig. 19]the threaded front end: record on skyrim's thread, replay on a second
 skyrim's thread                      replay thread
 ┌────────────────────────────┐
 │ eligible d3d11 call        │
 │  record into lock-free     │
 │  ring, 5 to 8 ns per call ─┼──→ replay: dxvk's real state
 │                            │    tracking, vulkan recording
 │ buffer writes return at    │
 │  once (allocation is       │    each buffer keeps two current
 │  already thread-safe)      │    pointers: one app, one replay
 └────────────────────────────┘
 deferred release. no atomics on the hot path

on 2026-09-23 at 15:20, campaign c21 ran the cpu-bound case (720p, low settings):

backendfps1% lows
native dx11345.0267
our fork (ring + threaded front end)354.7277
stock dxvk269.5

we beat native dx11 on the cpu by +2.8%, one day in, while stock dxvk trailed native by 22% in the same test. precision matters here, since the repeat that evening (c23) came out a tie inside the noise (-1.1%). the fair summary is “level with native on the cpu, one day in”. at ultra 1080p the story differed, because that scene is gpu-bound and native still led, 205.0 against 192.1 fps, -6.3%, in c23. the gpu track came next.

this front end also produced our best bug. roughly one run in ten crashed silently, starting from a build two days earlier, and a crash-time dump of the ring exposed the cause. when an allocation ended exactly on a lap boundary of the 4 MB ring, a packet’s write limit jumped a whole lap ahead. the packet grew past the end of the ring and wrapped. the replay thread walked into zeroed memory, read it as a constant-buffer rename of a null block, and crashed with 0xc0000005. a model of the allocator counted 56 boundary crossings per 4 million steps before the fix and zero after. in 32 loading-screen runs there were 5 crashes before the fix and none after.

the present path

next came latency, and a driver setting nvidia does not document turned it all around.

on the first night we noticed native skyrim presents through hardware independent flip, the fast path where the gpu flips straight to the screen. dxvk presented through “composed: copy with gpu gdi”, a slow path where the desktop compositor copies every frame. display latency for dxvk sat at about 25 ms, against native’s under 10.

nvidia’s driver offers a “Vulkan/OpenGL present method” setting that can present vulkan through the driver’s own dxgi swapchain (“layered” present). turning it on changed nothing. the driver detects dxvk and refuses to layer it, unless a hidden profile setting is also set: “Vulkan/OpenGL Present Method - Flags” (0x20324987). nvidia profile inspector lists a value for it, 0x00080004 (“allow promoting DXVK to DXGI/DirectFlip”), and that value unlocks it. campaigns c17 and c18 measured the outcome at display latency 24.3 ms to 7.9 ms, level with native’s 9.6 ms, on hardware independent flip.

that campaign also found a free gain in the compiler. the llvm-mingw (clang) build of the fork ran 3.4% faster than the msvc build of the identical source (154.0 against 148.9 fps). it became our default build.

bit 8 arrived after that. scanning the same flag bit by bit on top of the working value surfaced one more bit that mattered. 0x00080104 keeps the independent flip and cuts the gpu’s idle at frame start from 0.365 ms to 0.013 ms. in game that meant display latency 11.5 ms to 5.5 ms plus 1.8% fps, against native’s 10.0 ms. from then on our fork has put frames on the screen sooner than native dx11 in every campaign. its neighbour, bit 18, is a trap that looks similar yet forces the slow composed path (29.2 ms).

[fig. 20]the present path, step by step (display latency)
settingpresentation pathlatencynote
native dx11hardware independent flipunder 10 ms
dxvk, stockcomposed: copy with gpu gdiabout 25 ms
layered, flag offdriver refuses dxvkno change
0x00080004independent flip24.3 → 7.9 msc17, c18
0x00080104 (bit 8)independent flip11.5 → 5.5 ms in game+1.8% fps; idle at frame start 0.365 → 0.013 ms
bit 18forces composed path29.2 mstrap

bit 8 does not remove the cost of handing a frame to the display, it moves the cost elsewhere. campaign c48 measured every pass against native, and the whole frame came out only 0.15 to 0.18 ms heavier than native. 72 to 74% of that gap sat in one pass, the 4096x4096 sun cascade depth map. the present hand-off lands inside whichever gpu pass happens to be running at the time, and it most often takes the longest pass.

the way we measure latency matters. both numbers come from presentmon’s DisplayLatency; each backend is read on its own presentmon clock, which start at different points. on native the frame starts when skyrim’s own previous Present returns. on our fork it starts when nvidia’s layered swapchain presents on dxvk’s submit thread, after the frame was already replayed. presentmon’s GPULatency makes the difference apparent: 0.5 ms on our fork against 5.4 ms on native in whiterun. our number leaves out up to about two frames of queue. in whiterun, which is cpu-bound, the queues drain and “half” is plausible. in the traversal, which is gpu-bound, it is not proven. what holds is that the fork presents on the same fast hardware flip path as native.

we tried settling it on a shared clock. presentmon can time input to photon from the os input timestamp, so we tapped scroll lock every 250 ms during a capture. the taps never registered (0 of 6,231 frames), likely because skyrim reads its keyboard through directinput. a camera-based click-to-photon test is needed.

a correction, then ahead

two more levers closed most of the remaining gpu gap in whiterun.

the constant-buffer mirror. skyrim writes constant buffers to fast, cpu-cached ram that the gpu reads slowly. the obvious fix is placing the ring in vram with resizable bar. we built it and measured -30%. every locked cpu write to write-combined vram stalls about 200 ns while it drains over pcie, and skyrim performs about 13,000 of them per frame. the mirror keeps the cpu writes in cached ram. the gpu’s transfer queue then copies each used range into vram, in parallel with the graphics work. across 600 frames it moved 7.66 million renames with zero cpu reads and zero lookup misses, worth +2.8%.

vb rebar. dynamic vertex and index buffers get written in bulk and in order. write-combined vram handles that well // resizable bar is a gain for those buffers. a full-screen glare shader that read a 4 MB streaming buffer about 12.6 times per frame clued us to this angle.

campaign c53 then claimed a tie at ultra 1080p in whiterun, 219.0 fps against native’s 218.9. the 1% lows were 178 against 172, and presentmon display latency 5.1 ms against 9.6 ms.

an audit corrected it the same day. the median tied, but the mean of the five runs favoured native at 219.6 against 217.9. the perf-max profile already included half-rate volumetrics, so it was not vanilla’s exact image either. the claim and the correction stayed side by side in the log, and strict (visually close or identical to native) was defined because of it. under strict the fork ran 2.2% behind native.

hardware-accelerated gpu scheduling (hags) on helped every backend, but it helped us more: native +2.5%, strict +3.5 to 3.8%, perf-max +4.7 to 5.0%. acceleration off was worth trying but it stays on; hags off ended up decreasing frames for native and our fork.

campaign c58, on 2026-09-26 at 03:20, delivered the first paired success at ultra: +1.23% (95% ci +0.38 to +2.07, 8 of 10 pairs). latency was 5.2 ms against 9.5 ms. the same campaign surfaced a strange effect in dxvk’s descriptor buffers, which warm up across runs in one session: 205.5 → 209.4 → 216.9 → 215.9 → 211.4 → 202.6, then 219.2 → 216.0 → 218.1 → 219.2 → 219.0. a single warm-up run proved insufficient, and the standing rule is now roughly six runs of warm-up after any driver or dll change.

two exact shader replacements pushed further. the fork can swap a hand-written shader in for any of skyrim’s, matched by content hash. two blur shaders received bit-identical replacements (128 of 128 dispatches identical), worth +1.1%. vanilla’s volumetric-generate compute shader ran as 32x32 thread groups. the compute shader was rewritten as 8x8 with a matching dispatch and byte-identical output, plus a deliberately broken negative control to prove the check works. its dispatch fell from about 0.5 ms to 0.335 ms. with it, the whole perf-max stack reached +2.35% over native in whiterun, and the shader’s share of that was approximately +0.34%.

dead devices

for about a day, roughly one whiterun run in four lost the gpu device.

suspicion was on a feature called the getter shadow. OMGetRenderTargets and its siblings answered from an app-thread copy of the state instead of asking dxvk. its verification reported zero mismatches in 5.57 million slot comparisons. yet that verification also changed the synchronization it was checking, and it earned no measurable speed. our leading explanation is a lifetime bug: it answered from stale state after a view’s last reference had moved to the worker thread’s deferred release.

numbers spoke:

getter shadowwhiterun runs that lost the device
on13 of 47
off, same build0 of 12
before the feature existed0 of 57

the getter shadow was turned off by default and skips its record-time update when off // no device loss has returned since. the lifetime bug was never proven to be the cause, so we say the feature was implicated rather than the bug fixed.

along the way we had blamed one cluster of losses on a stale incremental build, a note that turned out wrong. those runs were bad luck at the base rate of about one in four, and the events were logged and examined. the detour did leave a rule behind: a merge that changes class layouts needs a clean build.

a second instability received the same treatment. async volumetrics (the volumetrics generate on a second graphics queue) ran slower in both profiles, -0.96% and -1.87%. the candidate build also caused four device losses and eight driver resets in under an hour. every driver reset appears as an nvlddmkm event 153 in the windows event log, checked after any risky run. async volumetrics was closed for good, keeping only its incidental bug fixes.

the traversal wall

whiterun fell first; traversal refused to move.

[fig. 21]the traversal scene: a no-clip flight across the plains outside whiterun
the traversal scene: a no-clip flight across the plains outside whiterun

the scene is gpu-bound. at about 4 ms per frame, the layered present hand-off’s fixed cost (about 0.36 ms) weighs far more than at 60 fps. for most of the project the traversal sat 5 to 9% behind native even while whiterun stayed comfortably ahead.

better instruments built for it:

  • a per-pass gpu timer wraps a vulkan timestamp query around every render pass.
  • a matching proxy dll does the same for native dx11, letting passes compare one to one without pix or nsight.
  • passdiff.py aligns two runs pass by pass and lists the biggest deltas.
  • a layout census counts image layout transitions and barriers per frame.
  • a “gaps” timer splits gpu idle time by cause.

those instruments returned an uncomfortable result. audit found that the per-pass floors already sum to about zero. in the best frame of each 120-frame window, our wins (the volumetric collapse, the blur replacements, the transfer-queue uploads) outweigh our fixed losses. the strict traversal gap of about 0.24 ms per frame fits the tail of the driver’s present hand-off. that tail lands in three passes of the following frame: the depth prepass, the far cascade and the volumetric generate. the 0.24 tail was incredibly difficult to isolate.

audit additionally landed two submission-structure leads:

  • zero-copy present (c67). the fork acquires the swapchain image early and points skyrim’s final draws straight at it, so the present blit disappears. worth +0.73% in the traversal, roughly the 0.03 ms the blit itself took.
  • flush at render-pass end (c70). dxvk flushes work to the gpu on internal heuristics without knowing whether a render pass is open. dxvk counted 86 to 88 render passes per frame where our pass timer counted 83. the difference was cutting implicit flushes in half. deferring those flushes to the next render-target change earned +1.48% in whiterun (+3.51% over native at that point).

and the list of things that measured null in the traversal kept growing:

  • a fourth swapchain image
  • a fix to the water cube’s shared depth
  • unified image layouts, where the census counted about 170 layout transitions per frame at no measurable price
  • sdma routing, whose counters read zero pending uploads
  • a fragment-shader framebuffer copy, worth -0.019 ms in a synthetic test and nothing in the game

what this left us, for this rig and five-pair campaigns, is that savings under about 0.03 ms rarely appear in the frame // changes to submission structure do, and so does temporal reuse. zero-copy sat at about 0.03 ms and did appear, just barely.

perf-max cracked traversal w temporal levers

we went where the frame time lives: work that repeats every frame without much need.

half-rate volumetric light (with a motion gate) ranked among the first perf-max levers, worth about +5.5% early on. the volumetric collapse folds vanilla’s 90 or so volumetric dispatches into one // bit-identical, from 0.41 ms to 0.14 ms.

half-rate water reflections (c73). skyrim renders a small cube map for water reflections every frame. the first version updated it every other frame. that delivered the first traversal win in the project: +1.23% over native (+0.40 to +2.07, 5 of 5 pairs). before we shipped, audit flagged that the reflection-target detection was too broad and had no motion gate. the hardened version, used from c76 on, narrows the detection and adds motion and age gates. it also adds a relaxed turn gate // a reflection cube does not change when the camera exclusivly turns.

the half-rate far cascade (c74). the far sun cascade (cascade 1 of the 4096x4096 atlas, ~2,280 draws in whiterun) re-renders every frame. on a reused frame our plugin skips it. the plugin also freezes that cascade’s camera, matrices and split data. that way the shadow mask and the volumetric generate read the depth with the same matrices it was rendered with. our first model of the cascade box failed because the engine scales it by 8, and a seat found and fixed that. unexpectedly this is a cpu gain // a reused frame skips ~2,280 caster registrations on the game thread.

sun tolerance (c75). while flying, the sun direction jitters by a hair every frame. at zero tolerance the cascade reused 48 of 7,202 frames, while a tolerance of 0.05 degrees reused 3,672 of 7,345, which is every other frame. the outcome was whiterun +11.26% and traversal +2.99%.

period 3 volumetrics (c76). the volumetric generate runs every third eligible frame instead of every second, adding up to 1.8% on top.

the c76 stack ended, and those are the numbers at the top of this report. all 48 of its runs passed every gate, with no device loss and no flagged contamination.

the ladder

the whole ladder in one place:

backend (whiterun, ultra 1080p)where it landed
stock dxvk 3.1.1161.8 fps in c80 against native’s 219.0: -26.13% in whiterun, -14.57% in the traversal, 0 of 5 pairs each
our fork, “perf” (cb ring + threaded front end)already +27.6% over stock dxvk by c23
our fork, strict (vanilla’s exact image)level with native in whiterun, about -5.5% in the traversal
our fork, perf-max+10.97% over native in whiterun, +4.37% in the traversal (c76). repeated in c80: +12.09% and +4.04%, which is 1.52x and 1.22x stock. lower presentmon display latency

mod fight

we finally got to the moonshot, not even a week later. the next question: could a player install several performance mods on native dx11 and dust us, and could our fork run those same mods?

both sides got the same mods, and we ran everything again.

display tweaks (c77) runs cleanly on our fork, and the fps lead holds. with it on both sides, whiterun reads +12.07% (+11.04 to +13.10) and the traversal +4.12% (+3.55 to +4.69). native’s median sat about 1.0% higher in whiterun and 0.5% higher in the traversal than in c76. that is a cross-campaign reading, not a paired display tweaks on/off run. the 1% lows narrow under it. display tweaks lifts native’s whiterun 1% low from 164 to 174 (ours 179), and in the flight the 1% lows are level at 217 against 215. display tweaks also enables dynamic havok timestep scaling (60 to 240 fps) on both sides, so the physics stick to each side’s own frame rate.

community shaders (c79) is the big one. it replaces most of skyrim’s lighting with its own shaders and adds effects like screen-space gi, skylighting and volumetric shadows. the price is ugly: native falls from about 218 to about 103 fps in whiterun. we ran the same mods on both sides, engine fixes (which it requires) and display tweaks (which uncaps native under it, see the traps below):

scenenative + csour fork + cspaired difference (95% ci)1% lows
whiterun103.1 fps120.3 fps+16.64% (+16.41 to +16.86), 5 of 5 pairs85.8 vs 114.4
flying traversal151.8 fps161.0 fps+6.05% (+5.52 to +6.57), 5 of 5 pairs117.8 vs 151.3

under community shaders the lead grows. its effects pile more load onto the cpu and the driver, where our constant-buffer ring and threaded front end help most. the images match, confirmed by our own screenshot-pair comparison and audits.

[fig. 22]native dx11 with community shaders, whiterun: 103.1 fps, 1% low 85.8, display latency 22.4 ms. display tweaks, community shaders 1.9 and engine fixes on both sides, median of 5 paired runs (campaign c79)
native dx11 with community shaders, whiterun: 103.1 fps, 1% low 85.8, display latency 22.4 ms. display tweaks, community shaders 1.9 and engine fixes on both sides, median of 5 paired runs (campaign c79)
[fig. 23]our fork with community shaders, whiterun: 120.3 fps, 1% low 114.4, display latency 10.0 ms. +16.6% fps over native (95% ci +16.4 to +16.9), 5 of 5 runs faster. each side's latency is on its own clock
our fork with community shaders, whiterun: 120.3 fps, 1% low 114.4, display latency 10.0 ms. +16.6% fps over native (95% ci +16.4 to +16.9), 5 of 5 runs faster. each side's latency is on its own clock
[fig. 24]native dx11 with community shaders, flying the plains: 151.8 fps, 1% low 117.8, display latency 19.6 ms
native dx11 with community shaders, flying the plains: 151.8 fps, 1% low 117.8, display latency 19.6 ms
[fig. 25]our fork with community shaders, flying the plains: 161.0 fps, 1% low 151.3, display latency 7.1 ms. +6.1% fps over native (95% ci +5.5 to +6.6), 5 of 5 runs faster
our fork with community shaders, flying the plains: 161.0 fps, 1% low 151.3, display latency 7.1 ms. +6.1% fps over native (95% ci +5.5 to +6.6), 5 of 5 runs faster

two notes:

  • our fork runs with fewer levers under cs. the volumetric collapse never engages because cs replaces the volumetric chain it looks for. the flush lever sits at its cap, though the far cascade reuse still works. native loses nothing from any of this, so it does not favour us. it means the gain comes from a smaller set of tricks than c76.
  • native’s 1% lows under cs carry hitches (worst frames of 18 to 20 ms against our 8 to 9 ms) in every repeat. we have not found the cause yet. driver-side shader compiles on native remain our suspect.

no grass in objects (c79), in its live mode, takes nothing measurable from either side. live mode means no cache, ray casting on, and the game’s own grass density. with it on both sides, whiterun reads +12.91% (+9.42 to +16.39) and the traversal +4.50% (+3.57 to +5.43), 5 of 5 pairs each.

two mishaps nearly turned into fake numbers:

  • community shaders was silently off. our first community shaders runs looked free, at native’s speed, because its log ended with Required DLL Data/SKSE/Plugins/EngineFixes.dll was missing. it loaded and then did nothing. with engine fixes staged, it came on and compiled its shaders before the main menu (3,577 files in its cache). native then ran at exactly 75.0 fps, this monitor’s refresh rate. community shaders rebuilds the swapchain (10-bit, “hardware composed”) without the tearing flag. native could not present faster than the refresh, while our fork presents its own way and was never capped. an uncapped fork against a capped native is no fair fight, so we voided those runs. display tweaks’ EnableTearing uncaps native, which is why the real test runs display tweaks, community shaders and engine fixes together on both sides.
  • the grass cache was sparser than the game. a downloaded grass cache for no grass in objects made native +13% faster in the flight. the screenshot explained why: most of the ground clutter and many bushes were gone. the cache had been built at a lower grass density than our settings, so we do not report it as a mod win // a cache built at our own density is on the list.
[fig. 26]native dx11 with the game's own grass: 239.8 fps, median of 5 runs (c76)
native dx11 with the game's own grass: 239.8 fps, median of 5 runs (c76)
[fig. 27]native dx11 with a downloaded grass cache (no grass in objects, nexus 78173): 271.8 fps, one run (c78s). the cache was built sparser than our settings, so most ground clutter and many bushes are gone. the +13% came from drawing less grass, so we did not count it. the fair test ran no grass in objects live, at the game's own density
native dx11 with a downloaded grass cache (no grass in objects, nexus 78173): 271.8 fps, one run (c78s). the cache was built sparser than our settings, so most ground clutter and many bushes are gone. the +13% came from drawing less grass, so we did not count it. the fair test ran no grass in objects live, at the game's own density

blursed

sometimes, we had to use some very normal methods. we kept a ledger of every blursed idea and how each one ended.

for the present path:

  • custom d3d12 present bridge. a second, minimal d3d12 device presents a zero-copy import of our vulkan image. it worked in a synthetic test (independent flip, idle down to 0.011 ms). but the d3d12 present still took about 0.4 ms wherever it landed, leaving the books even. one variant produced the best display latency we measured anywhere (8.0 ms at the time), without extra fps.
  • windows 11 composition swapchain. dwm stayed black until the buffers were also bound as shader resources, an obscure gotcha we found by bisection. it eventually rendered real content on an independent flip, but it paces at the refresh rate (75 to 77 fps). closed.
  • direct scanout, vr style (VK_NV_acquire_winrt_display). vulkan owns a monitor directly, as vr headsets do, but it needs a monitor removed from the desktop. we proved the extension exists and stopped there.
  • layered present off entirely. worth +4.5% fps at double the display latency. on a 75 Hz screen latency matters far more than frame time, so it stays off.
  • five hidden nvidia vulkan settings, swept across 14 values, with every row landing within 4 µs of baseline. closed in an afternoon.
  • fifo-first swapchain creation, a forum trick for the native path. no change on this driver.
  • “treat as native” flag values 0x803A5 and 0x802A5 (c63). neither helped, and dropping bit 8 lost about 1.5% while doubling latency.

for the gpu and cpu:

  • the cascade cache (redraw only moving shadow casters). at first it cached nothing, an unread c0 register held garbage and every mesh looked “moved”. fixed and exact, it skipped about 1,167 draws per frame and still lost 3 to 4%. the skipped draws were cheap depth-only draws, and about half the cascade (trees in the wind, actors) redraws anyway. parked.
  • the early mirror split lost the gpu device in 11 of 11 runs for a ceiling of 0.03 to 0.08 ms. parked.
  • variable rate shading. the adaptive mode lost 2.3%. the “constant 2x2” row in an earlier campaign had never run vrs at all. the config parser accepted only the literal string 2x2, and fell back to off without an error // audit caught it by reading the parser against the plan file.
  • water effects at half rate on the cpu. water took 9% of the main thread in our profile. the first version ran 25% slower because it called the engine’s update directly and paid a full update on skip frames // the fixed version was worth at most 1.7%, inside noise.
  • thread affinity. the windows scheduler beat every fixed-core layout we set by hand // giving smt siblings to the job workers lost 6%.
  • hidden engine ini switches (joblist active-wait, culling-plane optimization, front-to-back prepass, lod z-prepass) // none moved the frame rate.
  • reordering Main::Render to run the sun cull on a worker. reverse engineering found only 0x146 bytes of setup between the cull stage and the draw stage, and no independent work // it died on paper before making it to code.

shells

  • the profile inspector dialog ate a night. a cleanup script split a setting name on spaces and handed nvidia profile inspector a bad path. an elevated confirmation dialog then hung for about 9 hours (😭), from 04:21 to about 13:10. two overnight campaigns stalled until closed by hand.
  • game bar isn’t there. xbox game bar is not installed on this build. with game dvr on, windows kept trying to open an ms-gamingoverlay: link, and a popup stole focus about every 9 minutes. the harness refuses to type into a foreign window, so runs failed without explanation. the steam overlay was hooking Present on every run as well, and steam’s web helper leaked to 19.5 gb overnight, which stopped both queued chains (again, 😭). game dvr and the overlay are both off now.
  • a wait loop that waited for itself. a helper that waits for the bench lock turned into a busy spin during captures twice, for two different reasons. once, native read 170 fps instead of 196 to 215. on the last night, a new wait loop searched process command lines for “skybench.py”. it found its own command line and waited on itself for over an hour (… 😭).
  • engine fixes writes into its own ini. it fills EngineFixes_SNCT.ini with skyrim.esm values at runtime, so the harness’s hash check refused to remove the file we staged. the harness now names that file as game-written.
  • two cores spinning. with bJoblistActiveWait=1, skyrim’s six job workers burn two full cores at a still camera with nothing queued. switching it off changed nothing in the frame rate, and we did not prove the spinning actually stopped.
  • havok is nearly free in a still whiterun, with its workers using negligible cpu in our sampled still-camera run.
  • copy that is not provably dead. a late hdr copy looked dead and removable, and we traced it to the first-person render. then the water reflection cube’s faces turned out able to read it. it stays.
  • machine noise. twice, whole rounds of runs dipped 20 to 30% in every backend, stock dxvk included. both times the cause was a background job on the machine, not the code.

auditors

every major round of work was audited twice. both audits are read-only, working from committed evidence, campaign results and source at a named commit, and neither sees the other’s draft. eight rounds happened, 16 audit documents in all, and round eight was our last check.

each round paid for itself:

  • round one found that presentmon’s cpu and latency columns for the dxvk rows were timing the driver’s submit thread rather than skyrim’s game thread.
  • they found that the vrs “ceiling” row had run with vrs off.
  • they caught four device-lost runs that had passed the old gate and sat inside a campaign’s numbers // one of them had kept 2.3 seconds of frames and reported 228.7 fps. the harness got its device-loss fix from that audit.
  • both auditors, arriving from different directions, reached the same statistical floor. nothing under about 0.03 ms can be judged by a five-pair fps campaign on this rig.
  • on the strict traversal gap, one put the odds of any untried idea at about one in five. the other refused to call the tail a proven floor // the status line kept her hedge.
  • round eight recomputed c76 and c77 from the raw run files and reproduced both to the second decimal.

look ma i’m benchmaxxing

whatcount
benchmark runs with a result2,370 (2,144 ok, 226 failed and kept as evidence)
campaigns80
frames captured12,640,563
gameplay capturedabout 19.3 hours
benchmark data30 GB
busiest day587 runs (2026-09-23)
fork commits on top of dxvk 3.1.1300
fork diff245 files, +57,759 / -79 lines
fork experiment worktrees74
main repo commits437
skse plugin6,595 lines in 32 files
campaign plans72
seat briefs106
audit rounds8 rounds, 2 auditors, 16 documents
research notes43 files, about 133,000 words
tuning reportabout 18,400 words

every number in this report has its run id in the tuning report, in blessed-skyrim.

open items

  • the strict profile traversal gap. about 0.24 ms per frame, about 5.5%, sitting in the driver’s layered-present tail. the only route to close it is visibility work, such as conservative culling and gpu instance compaction // planned.
  • custom grass cache, built at our density, for a fair no grass in objects fight.
  • ideas for later: dynamic resolution, frame generation, and the ray tracer itself (stopped at gi v1.1 when we shifted immediate focus to performance.)
  • two machine angles not yet tested: gpu preemption, and large pages.
  • real click-to-photon latency test, with a camera, since injected key taps do not reach presentmon’s input tracking in skyrim.
  • native’s hitches under community shaders. cause not found, with driver-side shader compiles as the first suspect.

thank you

this project started when a very dear friend saw my minecraft work and asked me to optimize their favorite game (you know who you are), and they knew i would take this way too seriously. this project is the output of benchmarking nonstop for days, and despite starting off ~50fps under native, we didn’t stop and finally beat dx11 performance with vulkan :3

a week ago stock dxvk ran whiterun about a quarter slower than native, and it still does. today our fork runs it 11 to 12% faster than native, 1.5 times as fast as stock // it runs on the same fast flip path, with lower display latency on presentmon’s clock.

log

2026-09-28modsthe lead holds with the same mods on both sides (c77, c79). c80 repeats the headline. round eight audits the writeup
2026-09-27bothahead of native in both scenes, perf-max (c73, c74)
2026-09-26ultrafirst paired win at ultra, whiterun (c58)
2026-09-25levela tie at ultra in whiterun (c53). an auditor corrects it the same day
2026-09-24bit 8one more hidden nvidia bit. display latency under native's
2026-09-23cpuahead of native on the cpu (c21)
2026-09-23flagthe hidden nvidia flag: display latency 24.3 ms to 7.9 ms (c17, c18)
2026-09-23threadthe ray tracer leaves skyrim's thread (c10)
2026-09-23lightfirst light: ray-traced sun shadows in skyrim
2026-09-22sunthe sun's shadow mask, forced to zero from inside the fork
2026-09-22raysfirst hardware rays inside skyrim's vulkan device
2026-09-22bootskyrim boots on our own dxvk fork

report: closed. perf-max ahead of native in both scenes, twice. strict still about 5.5% short in flight. every number is attached to its run id.