Local AI PCs

Your own AI in the company. No cloud, no subscription, no data leaving the building.

We build computers that run large language models on your premises — contracts, accounting and source code never leave the company network. For every build we say which model fits and how fast it answers.

Why local AI instead of the cloud

Data stays with you

The model runs on your hardware. Prompts and documents go nowhere — for work under NDA, with personal data or medical records this is often the only viable route.

One-off investment

Cloud AI is billed per user per month. Your own machine serves the whole team with no query limits and typically pays for itself within a year.

Component prices are rising

Memory and graphics cards keep getting more expensive. A build bought today is also a hedge against the next wave — and the hardware keeps its value for other workloads.

Builds for local AI

For an AI PC the decisive number is not FPS but graphics memory. That is why every build lists its total VRAM and the models that will run on it.

Does your model fit?

Pick a model, quantisation and context length. Required memory is computed from model weights and KV cache — exactly, not estimated. Only the speed is an estimate.

Quantisation
Context length

Fits with headroom

Memory required10.5 GB
weights: 8.9 GB · context: 1.6 GBavailable: 21.9 GB

Out of 24 GB VRAM across 2 card(s), after driver overhead and the split between cards.

Estimated generation speed

~26 tokens per second (≈ words per second)

Rough estimate from the graphics memory bandwidth. Real speed depends on software and settings — measured values are added to the builds. Two cards add up memory, not speed — model layers run one after another.

Fits: the model runs with headroom even for longer documents

! Tight: runs, but with no room for longer context or a second model

Does not fit: needs more VRAM, or runs slowly from system RAM

Quantisation shrinks a model so it fits in memory: 4-bit (Q4) loses very little quality and is the standard for local deployment, 8-bit is a compromise, 16-bit is the uncompressed original.

Who we build this for

Companies that want to use AI on their own data and cannot or will not send it to the cloud.

  • Law firms and notaries — contracts and case files under confidentiality
  • Accounting and tax firms — invoices, statements and clients' personal data
  • Clinics and healthcare providers — records that cannot be shared
  • Development teams — a coding assistant without sending source to a third party
  • E-commerce and marketing — product copy and content at scale, no per-token fees
  • Manufacturing and service — search across manuals and internal knowledge bases

What it is good for — and what it is not

  • ✓ Running finished models (inference), an assistant over company documents, several users or models at once, fine-tuning smaller models with LoRA.
  • ✗ Full training of large models — that needs 48 GB+ VRAM per card and a fast link between cards. We will tell you when a local machine plus cloud makes more sense.
  • Two cards mean a bigger model and more concurrent users, not a twice-as-fast answer. We say so up front because it is easy to oversell.

How it works

01

Enquiry and proposal

Tell us what you want AI to do and how many people will use it. We recommend a build — or adjust the configuration to fit your model.

02

Build and test

We assemble the machine, burn it in and test it with the very models you will use. Speed is measured on the actual unit.

03

Handover and first run

On request we prepare the environment (drivers, Ollama or LM Studio, web UI) so the first model runs right after power-on. VAT invoice, 24-month warranty.

Glossary — what the numbers on AI builds mean

VRAM (graphics card memory)
The single number that decides whether a model runs at all. The model has to sit entirely in graphics memory; whatever does not fit runs from system RAM orders of magnitude slower. Two cards add their memory up — 2× 12 GB gives 24 GB.
Quantisation (Q4, Q8)
Shrinking the model so it fits in memory. 4-bit (Q4) is the standard for local deployment — it loses very little quality and makes the model more than three times smaller. A 14-billion-parameter model then takes about 9 GB instead of 28 GB.
Context
How much text the model holds in mind at once. 8k tokens is roughly 25 pages, 32k about 100 pages. Longer context costs more memory because of the KV cache added on top of the model.
Tokens per second
Answer speed. A token is roughly a syllable to a word; above 15 tokens/s the answer arrives faster than you can read it. The limit comes from graphics memory bandwidth, not from the processor.
RAG (answering over your own documents)
The approach where the model answers based on your contracts, manuals or policies. Documents are indexed and the relevant passages are handed to the model with the question — no retraining needed.
Ollama and LM Studio
Programs that run the model and expose it as a chat or an API for your own applications. They work offline, with no account and no subscription. On request we set them up on the machine.

Local AI PC — frequently asked questions

Which model will run on the build?
Total VRAM decides. With 24 GB you can run 4-bit quantised models up to roughly 30 billion parameters (Qwen2.5 14B, Gemma 3 27B, Mistral Small 24B). The 70B class needs 48 GB. The calculator above works it out exactly.
Why two graphics cards instead of one bigger card?
Two used cards with the same total memory are often far cheaper than one new large card. They add up memory and can serve two models or more users at once. They do not double the speed of a single answer.
How fast does the model answer?
For a quantised model up to 14B expect tens of tokens per second — faster than you read. For 30B models split across two cards, single digits to low tens. We measure exact values on every build.
Does our data really stay with us?
Yes. The model runs on your hardware; prompts and documents never leave your network. Tools like Ollama or LM Studio work fully offline, with no account and no subscription.
Is it loud, and how much power does it draw?
Under load this is a powerful machine — expect several hundred watts while generating and cooling you can hear. We offer a quieter office variant and a stronger server-room one; tell us in the enquiry where the machine will live.
Can the build be upgraded later?
Yes — every build states its headroom in power supply and slots. Typical upgrades are cards with more memory or a stronger CPU for parallel data processing.

Local AI PC enquiry

Tell us what you want AI to do, how many people will use it and where the machine will live. We reply within one business day with a recommendation.

Message us
Made with AI in Macaly