
LOCAL AI PC
AI PC – 2× RTX 3090 24 GB · 48 GB VRAM · Ryzen 9 9900X
48 GB of VRAM with both cards at PCIe 5.0 x8/x8 — runs even flagship 70B-class models.
Quick add-ons
Other
Microsoft Office
Extended Warranty
Two RTX 3090 cards add up to 48 GB of memory — enough to fit a flagship 70B-class model in 4-bit quantisation, meaning Llama 3.3 70B or Qwen2.5 72B. That is the answer quality you know from paid cloud services, except it runs in your own server room and nobody else sees your data. Smaller models such as Gemma 3 27B or Qwen3 32B run here with plenty of headroom and long context, so they can digest lengthy contracts in one go.
The ASUS ProArt X870E-Creator WiFi board is built for two graphics cards: both slots run at PCIe 5.0 x8/x8, full speed to each card. That is the difference from ordinary boards where the second card gets only four lanes — here tensor-parallel operation through vLLM makes sense, not just splitting the model across layers. On top of that: 10 Gb networking for moving models and datasets across the company network, four M.2 slots, dual USB4 and Wi-Fi 7.
Both cards are Turbo (blower) versions that exhaust hot air straight out of the case. With two cards stacked that is essential — ordinary triple-fan models would suffocate each other. The Fractal Design Torrent case is built for maximum airflow and the EVGA SuperNOVA 1300 W Gold supply has headroom even for the short power spikes the 3090 is known for.
A twelve-core Ryzen 9 9900X with 64 GB of DDR5 handles data preparation, document indexing for RAG and several concurrent users without holding the GPUs back. Two 2 TB SN850X drives keep models and company data apart. A PiKVM V4 Mini is included — remote keyboard and video level access, so you can reach the machine even when the operating system is down. For a machine in a server room or another branch that saves a trip.
What it is for: running large models for a whole team, an assistant over company documents, serving several users at once, fine-tuning smaller models with LoRA. RTX 3090 cards are also the only consumer models supporting NVLink — add a bridge and the two cards talk directly, off the bus. It is still not a machine for training models from scratch; that needs server cards with 80 GB of memory.
AI performance
48 GB
total VRAM
2× 24 GB
64 GB
RAM
for offload and data
2×
graphics cards
RTX 3090 · concurrent models
A build for flagship 70B-class models in 4-bit — quality comparable to cloud services, in your own server room.
What runs on this build (4-bit, 8k context)
| Model | Memory | Speed |
|---|---|---|
Qwen2.5 14B balanced business assistant, RAG over documents | Fits10.5 GB | ~69 tok/s |
gpt-oss 20B (OpenAI) OpenAI open-weight model, fast thanks to MoE | Fits13 GB | ~217 tok/s |
Mistral Small 3.1 24B European model, strong on documents and instructions | Fits15.7 GB | ~42 tok/s |
Gemma 3 27B strongest Gemma, text and images | Fits17.1 GB | ~37 tok/s |
Qwen3 30B-A3B (MoE) big-model knowledge at small-model speed | Fits19.1 GB | ~236 tok/s |
Qwen2.5 32B demanding writing and analysis | Fits21.8 GB | ~31 tok/s |
Llama 3.3 70B flagship class, needs 48 GB of VRAM | Does not fit45 GB | — |
Memory is computed from model weights and KV cache. Speed is a rough estimate from the graphics memory bandwidth; measured burn-in values are added over time.