Swift Made My RTX 5090 Feel Fast Again
From Huihui and ThinkingCap to Swift-1.5: the recipe changes, benchmarks, and everyday experience behind my latest RTX 5090 setup.
Sekou M. Doumbouya
Listen to this article
The views expressed here are my own and do not represent those of any current or former employer.
After we brought up Swift-1.5, I sent a message that was simpler than anything in the benchmark report: “This model seems really fast compared to what we were doing before when I use it.”
The responses were starting quickly enough that I noticed it while working. So far, my sessions have mostly been serial, with an occasional second sub-agent. It has been holding up. That is the experience I wanted from the RTX 5090.
Getting here took several rounds of testing. The questions kept changing: does the model work, can it handle enough context, what does that capacity cost in latency, and do my actual clients know what is running? Swift is the latest answer on my single-5090, 32 GB serving box. The experience is encouraging, and the data explains both why I like it and where I would be careful about adding more work.
Qualification Is Different From Usable Capacity
In the September 19 Huihui NInfer testing, the initial strict tool checks failed. I asked the agent to keep researching and try the appropriate settings. A compatible upstream runtime fix recovered the core and vision checks; that was a reason to continue qualification, rather than treat the first failure as a verdict on the checkpoint.
When we reached 64K, I pushed on the context limit: “64K is great for qualification, but is not enough for a usable context window.” This was a continuation of the usable-context problem I had already been writing about. The amount a server advertises and the amount an agent can use productively are different questions.
We expanded the Huihui configuration to 163,840 total tokens at concurrency one, with retrieval demonstrated at 150,144 input tokens and an 8,192-token output allowance. Roughly 4.5 GiB of GPU memory remained free. Estimates for 192K fell below the chosen 4 GiB reserve, so that larger profile did not advance. It was an estimate against our operating margin, not a failed 192K benchmark.
That is an important boundary. It was not a proven model maximum. It was a disciplined stopping point on this card, in this serve, with an explicit amount of room left for the machine to breathe. The context-envelope finding records the actual probe and the reason for the margin.
The next promoted recipe, ThinkingCap, made a different trade. We extended its own qualified 32K profile to 131,072 total tokens at concurrency one, retaining vision and speculative decoding while moving 3 GiB of weights into system RAM. On its strict short-output workload, median time to first token rose from about 0.46 seconds to 3.00 seconds. These were complete configurations with several differences, so that comparison does not isolate the cost of offload alone. It does show the responsiveness we gave up to retain the larger window in that recipe.
For the next experiment, I wanted to change the order. “I don’t want to go and ramp up all the individual settings to get to the end,” I told the agent. I had found a complete recipe reported to work on similar hardware. I wanted to start there, get it running with the functional checks, and then benchmark the configuration I intended to use. This was a choice about this experiment, not a claim that incremental tuning has no value.
The Swift Recipe Was a Complete Starting Point
The starting point was a single-RTX-5090 report and its Swift-1.5 Qwen3.8-27B NVFP4 artifact. We pinned and tested that artifact. I also shared a separate uncensored MTP checkpoint, but it remained a research lead; these results are not from that model.
The retained profile has 262,144 total context tokens, K8V4 KV cache, DFlash2 with seven draft tokens, vision enabled, a 16K default thinking budget, concurrency four configured, a 32,768-token output limit, and a 450 W power cap. Weights occupy roughly 18 GiB of VRAM. Input and output share the total context window.
The adaptation was mostly outside the GPU. The card’s launch recipe requested 48 GiB of pinned host KV cache; this Windows 11 / WSL2 host has 30.91 GiB of physical RAM. We reduced host KV to 2 GiB and host state slots from 16 to 4. Same GPU model does not mean the rest of the machine is interchangeable.
I also removed --image-token-budget 1280 after startup rejected it. Vision did not disappear with that rejected flag. The resulting recipe passed 12/12 OCR and image checks. That detail is mundane, but it is exactly the kind of thing that gets lost when a recipe is copied as a block of flags and called done.
The publisher reports 160.8 tokens per second on a 256-token probe, 3,269 prefill tokens per second at 200K, plus IFBench and GSM8K results. I did not reproduce those claims locally, so I will not repeat them as my results. My serve showed 282,112 shared-GPU KV tokens, below the publisher’s 308,736. I have not isolated the cause of that gap, and I will not pretend that one host-KV change proves why it exists. The promotion record carries the configuration and evidence.
Speed Changed the Feel of the Work
My impression of responsiveness is stronger than my ability to assign it a clean speedup ratio. We did not run a matched old-versus-new time-to-first-token experiment. The previous ThinkingCap capacity tests required exact 32-word answers; the Swift runs below used a different output target and validation mode.
The capacity probes give that feeling some shape. They used warm serves, unique prefixes and canaries, temperature zero, thinking disabled by request, a 256-word target, and a 512-token cap. Zero outputs were exactly 256 words, so these are descriptive runs, not controlled output-rate measurements.
| Prompt size | Concurrency | Completed | Mean visible TTFT | Decode / aggregate observation |
|---|---|---|---|---|
| 3.7K | 1 | 100 / 100 | 0.31 s | 465 tok/s decode |
| 3.7K | 4 | 100 / 100 | 0.43 s | 195 tok/s per request, 615 tok/s aggregate |
| 31K | 1 | 30 / 30 | 3.04 s | Capacity probe completed |
| 31K | 4 | 30 / 30 | 3.83 s | Capacity probe completed |
| 126K | 1 | 10 / 10 | 24.75 s | Capacity probe completed |
| 126K | 4 | 5 / 10 | 44.85 s, successes only | Five queue timeouts near 30 s |
| 252,223 | 1 | 3 / 3 | 82.09 s | Capacity probe completed |
Choose a measured prompt size and concurrency. These are recorded observations, not a capacity forecast.
All requests returned visible output. Output lengths varied.
All recorded cells
| Workload | Completed | Mean visible TTFT |
|---|---|---|
| ~3.7K prompt / 1 active request | 100 / 100 | 0.31 s |
| ~3.7K prompt / 4 active requests | 100 / 100 | 0.43 s |
| ~31K prompt / 1 active request | 30 / 30 | 3.04 s |
| ~31K prompt / 4 active requests | 30 / 30 | 3.83 s |
| ~126K prompt / 1 active request | 10 / 10 | 24.75 s |
| ~126K prompt / 4 active requests | 5 / 10 | 44.85 s* |
| 252,223 prompt tokens / 1 active request | 3 / 3 | 82.09 s |
Those 465 and 615 figures should not become a claim that I beat the publisher’s 160.8. Per-request decode and aggregate throughput measure different things, and our prompts, output lengths, and thinking behavior are not a matched reproduction.
Short prompts held up at concurrency four. At 31K, all requests still completed in this test set. At 126K, five of ten requests hit HTTP 503 queue timeouts. No GPU out-of-memory event or container crash was observed. Four 126K prompts alone exceed the shared 282K KV pool, before output reservations. The configured concurrency limit does not provide four full context windows.
So far, my own sessions have been serial. At most, I might spin off a second sub-agent to do something. The serve has been holding up for that. That is encouraging, but it is not dedicated concurrency-two qualification, and it definitely does not prove a four-agent workload at maximum context.
Quality Still Needs Its Own Evidence
Large context only matters if the model can use it. I ran routed retrieval through 257,898 actual input tokens with 4,096 output tokens reserved and got 18/18 successes. That is a needle-style retrieval check, not a broad reasoning benchmark. The synthetic agentic suite reached 17/18; the one failure was a protocol failure despite a correct final answer. A frozen, officially graded SWE-Verified set completed 5/5 issues, which is evidence for those five issues and nothing more. Image validation remained 12/12.
I have not reproduced IFBench or GSM8K, established broad quality superiority, or completed a long production soak. The default thinking behavior also differs from the capacity probes. My engineering read is that Swift is promising for the way I am using this card: interactive work with limited overlap. The quality checks support specific capabilities; they do not settle every coding or reasoning comparison.
Operations Are Part of the Recipe
I was not willing to call the change complete just because a benchmark script had green output. I wanted to know whether the pieces around the serve had moved with it. That meant synchronizing the Pi, OpenClaw, Hermes, and WebUI model catalogs; updating Grafana panels; correcting a stale Pi display name to Anvil llm.secondary; restarting the Pi Web surface; and verifying the fresh catalog and a tool smoke. A repeat deployment reported zero changes.
There is one limitation still open: an authenticated visual check of the Workbench browser UI remained pending. I will not say every browser surface accepted the change until I have that evidence. Infrastructure is a relay race. A fast server does not help much if the client still presents yesterday’s model.
The principle I am taking forward is simple: measure the window you intend to work in, under the pressure you expect to create. A local model earns its place when the full path holds together: the recipe starts, the quality checks are honest, the clients point to it, and the context remains usable. The hardware will change. That discipline should not.
Co-authored with AI, based on the author's working sessions, dictations, and notes.
Explore the source
fakoli/anvil-serving
This article discusses an open-source project. Star it, fork it, or open an issue.