A Card That Made Me Sign a Letter
Why I added a second RTX PRO 6000, upgraded my CPU, and then watched DeepSeek 0731 land on the exact card I just bought.
Sekou M. Doumbouya
Listen to this article
The views expressed here are my own and do not represent those of any current or former employer.
I went to buy a graphics card for my own home lab, and the vendor asked me to sign a legal document first. Not for warranty. Not for payment. A letter shielding them in case something bad ever came of how the card was used, part of the export rules now settling over these parts. It made sense. It also felt bizarre. I do not customarily sign agreements just to purchase a computer part.
The disorientation was the point. I had been on the fence for months, and holding that letter in my hand forced the question I had been avoiding: what am I actually doing this for, and to what end.
That letter is the reason this post exists. It was the moment what I am building stopped being a hobby and became a commitment. So I want to walk through the whole decision, the topology, and the discipline that had to follow, because the card was never the hard part. The hard part was everything I had to believe to buy it.
A Bet I Laid Down on My Own Desk
The question that started this, the one I kept circling, was not really about which card was faster. It was about what kind of machine this was going to be.
I had been running one RTX PRO 6000 and a 5090. On paper that looked generous: two distinct lanes, a 96 GB professional card for the serious model work and a fast 32 GB consumer card for gaming and quick side services. It was a capable hybrid, and it quietly capped me. Some models were out of reach, and the ones I could run never quite felt like we were going anywhere.
The mismatch was the problem. Different VRAM, no NVLink, so they ran as independent replicas or wasted the smaller one. Neither was the move I wanted. Duplicating a model across two uneven cards is a way to double your power bill without doubling your capability. I wanted symmetry: two cards that tensor-parallel properly, so the same model runs larger, faster, and stays resident on hardware I own.
So the decision became a change in identity, not a number going up. I was not swapping a 32 GB card for a 96 GB card. I was turning a mixed 96 plus 32 platform into a symmetric 96 plus 96 platform. Hybrid workstation out. Model lab in.
Toggle the box. This is the whole decision in one picture: a capable hybrid doing two jobs, or a pair of matched cards committed to one.
- Capacity lost on the smaller card when the big one is busy.
- Models sharded across uneven cards, or run as wasteful replicas.
I will be honest about why I wavered. The 5090 is actually faster than the RTX 6000, and giving it up, even at 32 GB, felt like a loss. I liked gaming on it, and the 6000 is not a gaming card. That concern melted once I realized the PC I had gutted to build this one still had enough parts to hold the 5090. I did not lose anything. Today the 5090 lives in its own box, still doing gaming and model experiments.
The irony is not lost on me: I signed a legal document to buy the quietest possible version of the loudest thing I own. The Max-Q runs cool and silent because I spent years in data centers, and I do not miss the hum of rack servers in my living space. The machine doing the heavy lifting should not sound like it is.
The Topology Behind the Fence
The second card was only half the story. The rest was rearranging what I already had so it could earn its place.
First, the display had to move. The RTX 6000 Max-Q draws a hundred watts less when it is not driving a monitor, so I moved video output onto the CPU’s integrated graphics. That freed both cards for pure inference. It sounds small. It was not. It was the difference between a card doing one job and a card doing the job I bought it for.
Second, the CPU had to change. I had been hitting a wall compiling Triton and mixture-of-expert kernels. The 9800X3D, with its eight cores, could not keep up. I upgraded to a 9950X3D2, and the difference in compile throughput was real, especially for MoE kernels. You cannot see that on a model benchmark, but you feel it every time you build. The model server is only as fast as the box that has to keep it fed.
The host RAM was the decision I deliberately did not make. More system memory would have been nice alongside all this GPU memory, but my next logical step is a Threadripper board, and that socket would make whatever kit I bought today useless. So I kept 96 GB of high-speed memory and overclocked it at EXPO. Buying memory I would throw away at the next socket is not a win, it is a lease with extra steps.
Third, I had to accept what the pair is not. Two cards tensor-parallel over PCIe, with no NVLink between them, is not a single monolithic slab of VRAM. Both enumerate at PCIe x8. That would bother me chasing frame rates, but for large inference, the binding constraint is memory, not lane purity. The second card is capacity first, concurrency second, speed third. Matching that reality meant choosing the right serving profile instead of pretending the hardware was something it was not. I will come back to that, because it is the point of the whole post.
Here is the honest version of what the pair unlocks, because “more VRAM” hides the real shape. The 120B tier, Qwen3.5 122B at NVFP4, already fit comfortably inside a single 96 GB card. So the second card was never about opening a door that was closed. It was about how models get to run: at higher precision, with more room for context and KV cache, resident for longer, and across a computed pair instead of a compromise. And it opens the only path toward the 200B-plus MoE class, the models that would not fit a single 96 GB card even at aggressive quantization, the class that DeepSeek V4 Flash 0731 turned out to be.
One 96 GB card already runs the 120B class. The second card is not about access to that tier. Toggle the box and watch which model classes fit, and how.
Measured model memory (GiB)
| Model | Model mem | KV cache | Where |
|---|---|---|---|
| Qwen3.5 122B A10B (NVFP4) | 73.22 | 13.84 | one card |
| Laguna S 2.1 (NVFP4) | TP=2 qualify | two cards | |
| Nemotron 3 Super 120B (NVFP4) | TP=2 + EP=2 | two cards | |
| DeepSeek V4 Flash 0731 | 650K ctx, TP=2 | two cards | |
Power and noise shaped the card choice as much as memory did. After years in data centers I did not want rack noise in my home, and the Max-Q runs cool and silent. It was a no-brainer: less power, less noise, still 96 GB.
Why I Bother Owning It
Let me be blunt. You do not spend on owned hardware hoping it will never be outpaced. Frontier labs are racing cost down every quarter, and I will not win that race. This is not a hedge against a future I expect to lose. It is a position I am choosing to hold while the technology is merging around me.
The reasons to own are the ones that do not show up on a price tag:
| Reason | What it actually buys |
|---|---|
| Privacy | My data stays on a box I control. I am not renting from somewhere I do not fully see. |
| Architecture literacy | Running many models myself, feeling the differences between their architectures, not just reading about them. |
| Independence | No RunPod, no VAST, no rented instance where I cannot really know where my data goes or who can see it. |
| Room to push | Taking a model to its extreme, not just to its most comfortable point on a rental. |
But I want to be honest about the cost of this table, because that is the part people skip. Owning is cheaper per token at the extreme and more expensive everywhere else. It puts the maintenance, the cooling, the power, the upgrades, and the failures all in my lap. A rental fails and you file a ticket. My box fails and it is mine to fix. The reason the table still wins for me is that the failure mode I care about most, the one where I lose the ability to even try, does not appear on a rental invoice. Independence is not a discount. It is a different kind of bill.
Then DeepSeek Just Showed Up
Here is the part that felt like a coincidence. I bought the second RTX PRO 6000 for reasons that had nothing to do with anyone’s release calendar. And on a Friday at the end of July, DeepSeek dropped V4 Flash 0731, the official release of their Flash model, and it happened to be a model that fit this exact card. Not barely. Comfortably. It was the Friday surprise that made the whole decision feel less like a hedge and more like scaffolding.
The timing mattered, and I say that without claiming I predicted it. I did not. But the model and the hardware were speaking the same language, and that is worth pausing on.
DeepSeek V4 Flash 0731 is a 284-billion-parameter mixture-of-experts model with 13 billion active per token. That is exactly the class the second card exists for. It does not fit one 96 GB card. It clears the bar of the pair, and it lands low enough in activated parameters that inference on two 96 GB cards is genuinely plausible. Same architecture as the earlier preview, same size, same 1M-token design. The jump is in the post-training: agentic and coding gains so large that this Flash release beats the earlier Pro preview on every coding and agent benchmark DeepSeek published, at a fraction of the cost. It ships MIT-licensed and ungated, so running it myself, commercially, on my own substrate, is not a legal gray area. That matters to someone who just signed a letter to buy the card.
I want to be honest about what those numbers are worth. They are vendor-reported, run on a harness DeepSeek has not released, and rated on internal test sets. That is a reason to test it on my own traffic, not a reason to trust a model card. What I can say from my own box is narrower and more telling: it competed with models I had been defaulting to, and I preferred using it. It brought real life back into the harnesses I had built. I spent more time with agents I had mostly left running, and using them taught me more than reading about them ever would. That is the part a benchmark cannot show you.
The coincidence resolved into a conclusion. I did not buy the card because DeepSeek was about to release something perfect for it. I bought it to have room to push, and a model showed up that rewarded exactly that room.
The Discipline Is Where the Leverage Lives
None of this works if a benchmark talks me into shipping something real traffic cannot survive. The card gives me room to push. The discipline tells me when to stop pushing.
The promoted 650K profile measured 141.6 output tokens per second at 32K context, about 8,800 prefill tokens per second, and a needle retrieval past 640K that passed. That is the number that shipped. Here is the full gate:
| Envelope | What the gate showed | What I shipped |
|---|---|---|
| 128K | baseline, stable | kept |
| 650K | ~640K retrieval, real client smokes passed | promoted as llm.primary |
| 1M | ~985K retrieval, synthetic gates passed | rejected |
Now read that table the way an engineer should, not the way a benchmark wants to be read. The 1M envelope passed its synthetic gates. It even recovered a needle near 985K. Then real coding traffic crashed it. Twice. Two crashes under actual client shapes, not under a synthetic harness. A configuration that wins in a benchmark and loses in production is not a win. It is a trap with good numbers on the front.
So I promoted the 650K that survived real client shapes, and I kept the 1M evidence in the record instead of deleting it. The failure is part of the story. The crash tells me exactly where the ceiling is, and that is information I would rather keep than paper over.
Here is the trade that sets the ceiling. The model server splits its budget between two things you can turn against each other: how many sequences it admits, and how many tokens it batches. Holding the model, image, and topology fixed, you can grow the context window only by shrinking the batching budget. The slider shows what that did on my two-card box.
Same two cards, same model, same image, same quantization. Only the batching budget moves. Slide it and watch what context window becomes reachable.
Illustrative shape from measured envelopes (650K at 4096 batched, 1M at 2048 batched). Real reach on your hardware depends on card memory, quantization, image, and your own ceilings. The point is the trade exists: growing the window costs batching throughput.
A Long Bet, Taken While the Cycle Turns
I do not know if these cards halve in price next year. That is not the bet. The bet is that the knowledge I am earning right now, while the technology is merging and I am inside that cycle, is worth it if you can afford the entry. It is not a hedge against a future I expect to lose. It is a decision to be part of the merge instead of watching it.
This is where I stay honest, because the 650K ceiling was a narrow win. It ran with about 94 MiB of reserve headroom. That is not a margin I would bet a co-residency strategy on. It is a margin I would measure again on a different host before I generalized anything. The matched pair is a lever, not a promise. It points in a direction, and it still has to be operated honestly every day.
I came into this asking what the largest model I could run was. I came out of it asking something else: what kind of local AI platform can I build? The second card answered the second question. It turned a powerful desktop with AI bolted on into infrastructure, a machine that can host models, compare them, route work between them, and keep several ready at once. That is why I own it. Not for one benchmark, but for the posture it buys.
The anvil stays. The hammers rotate. The substrate I own lets me stay in the work, and the work is what compounds. Not the hardware. Not the model. The work.
Co-authored with AI, based on the author's working sessions, dictations, and notes.
Explore the source
fakoli/anvil-serving
This article discusses an open-source project. Star it, fork it, or open an issue.