LOCAL AI PC
AI PC – 2× RTX 3060 12 GB · 24 GB VRAM · 64 GB RAM
24 GB of VRAM across two cards: a company document assistant up to 30B parameters. Data stays with you.
Quick add-ons
Other
Microsoft Office
Extended Warranty
A computer built for one job: running large language models inside the company. Two RTX 3060 cards add up to 24 GB of memory, enough for 4-bit quantised models up to roughly 30 billion parameters — Qwen2.5 14B and 32B, Gemma 3 27B, Mistral Small 24B or OpenAI's gpt-oss 20B. The second card also lets you run two models at once: an assistant for the team and a second one processing documents in the background.
64 GB of RAM handles offloading of larger models and data work; the 2 TB WD Black SN850X NVMe loads a 20 GB model in seconds. The 1200 W Gold power supply leaves headroom for swapping in stronger cards later. Three industrial Noctua fans keep temperatures in check through hours of continuous load — this is a machine for a server room or utility room, not for a desk next to people. A quieter cooling variant is available for offices.
What it is for: running finished models (Ollama, LM Studio, vLLM), an assistant over company documents (RAG), several concurrent users, fine-tuning smaller models with LoRA. What it is not for: full training of large models — that needs 48 GB cards and a fast interconnect. The second card sits in a PCIe 4.0 x4 slot; for inference that does not matter since weights load once, but it would bottleneck tensor-parallel training. And two cards add up memory, not the speed of a single answer.
For an extra CZK 5,000 we swap the Ryzen 5 5600 for the sixteen-core Ryzen 9 5950X. Worth it if you will process data in parallel, serve more users or fine-tune; it does not speed up the answer generation itself, which runs on the GPUs.
AI performance
24 GB
total VRAM
2× 12 GB
64 GB
RAM
for offload and data
2×
graphics cards
RTX 3060 · concurrent models
A build for models up to ~30 billion parameters in 4-bit — the class most company document assistants run in today.
What runs on this build (4-bit, 8k context)
| Model | Memory | Speed |
|---|---|---|
Qwen2.5 14B balanced business assistant, RAG over documents | Fits10.5 GB | ~26 tok/s |
gpt-oss 20B (OpenAI) OpenAI open-weight model, fast thanks to MoE | Fits13 GB | ~83 tok/s |
Mistral Small 3.1 24B European model, strong on documents and instructions | Fits15.7 GB | ~16 tok/s |
Gemma 3 27B strongest Gemma, text and images | Fits17.1 GB | ~14 tok/s |
Qwen3 30B-A3B (MoE) big-model knowledge at small-model speed | Fits19.1 GB | ~91 tok/s |
Qwen2.5 32B demanding writing and analysis | Tight21.8 GB | ~12 tok/s |
Llama 3.3 70B flagship class, needs 48 GB of VRAM | Does not fit45 GB | — |
Memory is computed from model weights and KV cache. Speed is a rough estimate from the graphics memory bandwidth; measured burn-in values are added over time.