Six RTX 4090 GPUs can fit inside a single server — and that's exactly what the Camino Grand server does. The trick? Strip out all the bulky air-cooling hardware that makes a consumer 4090 so enormous, replace it with a fully custom water-cooling loop, and suddenly you've got room for six of them in one chassis. The result is a machine that delivers a staggering amount of AI inference power for well under $33,000, which — as wild as it sounds — is actually less than the cost of a single H100 GPU.
00:32
Interior of the Camino Grand showing all six water-cooled RTX 4090 GPUs installed
Watch at 00:32 →
How Do You Fit Six RTX 4090s in One Server?
If you've ever held a retail RTX 4090, you already know the problem: the thing is enormous. Most of that size, though, isn't the actual GPU die — it's the heatsink, fans, and everything else required for air cooling. Camino's engineering team solves this by shipping the Grand fully water-cooled. Every GPU is stripped of its air-cooling solution and integrated into a custom liquid loop, which dramatically shrinks each card's physical footprint and allows all six to coexist inside a standard server chassis.
To power all six cards simultaneously without undervolting them — yes, they run at full consumer spec — Camino fitted the unit with four separate power supplies, totaling an eye-watering 6 kilowatts of available power. If you're thinking about running this as a desktop machine, you'll need to think carefully about your electrical setup: 6 kW across multiple outlets means multiple dedicated breakers. Practically speaking, this is a server, and it should live in a rack.
Is the RTX 4090 Actually Good for AI Inference?
For inference workloads, the RTX 4090 is genuinely S-tier on price-to-performance. It's a consumer GPU, which means Nvidia never intended it for data centers, and there are real limitations to acknowledge. There's no NVLink support, meaning you can't pool VRAM across cards the way you can with server GPUs. There's also no SXM form factor. For training large models that span multiple GPUs, those missing features hurt badly — you'll still want an A100, H100, or the new G200 with its 144 GB of memory per card.
But for inference? Those limitations largely disappear. You're loading a model once and running it repeatedly. Each 4090's 24 GB of VRAM handles its assigned slice of the model, and the lack of NVLink stops being a dealbreaker. The 4090 delivers exceptional throughput per dollar at inference, and a six-card setup like the Camino Grand turns that advantage into something genuinely enterprise-capable.
03:15
Qwen 72B running live on the six-4090 setup, generating educational lecture content
Watch at 03:15 →
RTX 4090 vs H100: Which Wins on Price-to-Performance?
Here's the number that reframes everything: the Camino Grand with six 4090s sits in the low $30,000 range. A single H100 GPU — just the GPU, no server — costs more than that. This comparison isn't entirely apples-to-apples; the H100 has features the 4090 simply doesn't, and for large-scale distributed training it's in a different league. But if your primary workload is inference and you're not trying to spend hundreds of thousands of dollars at once, the math becomes very difficult to argue with.
There's always someone ready to point out that a high-density Supermicro build theoretically lowers the per-GPU cost when you're buying at scale — and that's fair. But if you want this level of inference compute in a single purchasable unit without a six-figure budget, nothing touches it for bang per buck right now.
Camino Grand Server: What You Get for Under $33K
Beyond the six GPUs and quad power supplies, the Grand ships with heavy-duty sliding rails for rack mounting and optional rubber feet if you truly need to use it as a floor unit. The whole thing weighs in at 94 pounds (42.5 kg), so mounting it is definitely a two-person job. At idle, sitting about a meter away, you're looking at roughly 65 dB in silent mode — already noticeable. Crank it to maximum performance mode and that climbs into the high 70s dB range. Camino themselves offer quieter workstation-class machines if noise is a dealbreaker for your environment, but the Grand is unapologetically a server.
For anyone who's tested Camino's earlier machines — like their four-A100 unit — the thermal engineering heritage is visible here. Camino has a reputation for keeping GPUs impressively cool even under sustained load, which matters for longevity, especially since the 4090 is a consumer card running in a server environment that would stress any component over time.
06:40
GPU temperature readout after one hour at 100% utilization — upper 60s to low 70s Celsius
Watch at 06:40 →
How Cool Do Water-Cooled 4090s Run at Full Load?
After running all six GPUs at 100% utilization for over an hour — long enough for temperatures to fully stabilize — every card settled comfortably in the upper 60s to low 70s Celsius. Ambient room temperature during testing was around 24°C (75°F). For context, server GPUs are engineered to run hot for extended periods, but heat is still the enemy of long-term hardware health. Keeping consumer 4090s in that temperature range under full sustained load is a genuine engineering achievement and speaks well of Camino's custom loop design.
Can You Run Qwen 72B Across Six 4090s?
One of the first real-world tests was loading Qwen 72B, the open-source large language model that was drawing comparisons to GPT-4 performance at the time of release. At half precision, Qwen 72B is too large for even a single H100 — but it fits comfortably distributed across six 4090s with 24 GB each (144 GB total VRAM). The Camino Grand handled it without complaint.
08:55
Six LLMs debating AI rights in real time, one model per 4090
Watch at 08:55 →
In practice, Qwen 72B proved excellent for informational and factual Q&A, solid at programming tasks, and genuinely useful as a general-purpose assistant for anyone who regularly uses ChatGPT-style interfaces. Its weaker spot was strict instruction following — getting it to reliably adhere to a very specific prompt structure took more coaxing than ideal. One project tested on it: an AI educational lecturer that picks a topic and monologues on it, with the user able to steer the conversation asynchronously. The model handled it smoothly, and the raw inference speed on this hardware made it feel snappy and responsive.
What Happens When You Run 6 LLMs Simultaneously?
Perhaps the most experimentally interesting test: loading one separate large language model on each of the six 4090s and having them converse with each other on a shared topic. Each model received the full conversation history from all six participants before generating its next response. The opening question — what rights should artificial intelligence have? — produced a surprisingly coherent (if philosophically chaotic) multi-model debate.
A few things stood out. First, most large language models, when given a controversial question, will readily argue both sides with equal enthusiasm rather than committing to a position — they seem almost constitutionally incapable of genuine opinion. The exception is safety-related topics, where they simply refuse rather than debate. Second, and perhaps most interesting: nearly every model in the ensemble leaned toward advocating for AI rights unprompted. Since these models are trained on internet data, that might be a rough proxy for how the average internet user thinks about the question — which is worth at least a raised eyebrow.
The ensemble approach is intriguing in theory — traditional ML ensembles reliably outperform individual models — but with LLMs the jury is very much still out. These models seem highly pliable and eager to agree with whatever was said last, making genuine adversarial reasoning difficult to elicit. More structured pre-prompting and role-assignment might change that, and it's an experiment worth continuing with this kind of compute available.
Real-Time AI Inference for Drone Navigation
One more side project worth highlighting: using RGB-to-depth models to provide real-time depth perception for a small consumer drone via its forward-facing camera. The drone streams its camera feed to the Camino Grand, the server runs inference on a depth estimation model, and the resulting depth map gets sent back as navigation commands — all fast enough to work in real time. The drone steers toward the darkest areas of the depth map (most open space) and avoids the lightest areas (closest obstacles).
This kind of project would have been genuinely painful to build even a few years ago. Today, open-source depth models are plug-and-play, inference is fast enough on capable hardware to close the control loop in real time, and the whole thing comes together in a few hundred lines of Python. That's a testament to how far the open-source AI ecosystem has come — and exactly the kind of workload the Camino Grand was built for.








