← build log Build log

Running a private AI assistant on a Raspberry Pi 5

JACOB DECAMP · FIRST PUBLISHED AUGUST 2026 · UPDATED SEPTEMBER 30, 2026

PHNTM One is a Raspberry Pi 5 AI assistant that runs its conversation model, memory, document retrieval and voice on the device. In its default Private mode, your conversation is processed locally. Optional cloud modes are separate choices and send data to the services you enable.

The engineering challenge is fitting all of that beside a touchscreen interface in 8 GB of RAM. Here is the hardware, the software stack, the failures that shaped it, and the evidence you can inspect before reserving one.

Explore before you build or buy

Try the interactive PHNTM One demo for the product experience, or use our offline AI guide to compare a desktop runner, a DIY build and an appliance. The browser demo simulates the interface.

The Raspberry Pi AI assistant hardware

  • Compute: Raspberry Pi 5, 8 GB RAM, quad-core Cortex-A76.
  • Interface: 10.1-inch 1920 × 1200 touchscreen, built-in speaker, USB microphone and mini wireless keyboard.
  • Storage and power: 256 GB microSD in the listed production configuration, active cooling and a 27 W USB-C supply. NVMe is an upgrade path; some prototype measurements used it.

The full specifications distinguish the supplied configuration from upgrades. The bill of materials lists the parts and their costs.

Local model, memory and voice stack

The published stack uses Debian 13, Ollama / llama.cpp for inference, Gemma 3 4B with a Llama 3.2 3B fallback, and nomic-embed-text for memory and document retrieval. whisper.cpp transcribes a tap-to-talk request locally; Piper generates the spoken reply. Voice processing shares the same memory budget as chat.

A model runner is one component. A usable assistant also needs document ingestion, remembered context, a clear indication of the answering model, and recovery when a worker stops doing useful work. Those are the parts that make this an appliance.

The budget

8 GB sounds like plenty until you write the ledger down. The OS, the display stack, and the assistant's own services want their gigabyte. The quantized 4B model (Gemma 3 4B, Q4_K_M) has model files of a few gigabytes, but its runtime memory also includes context buffers and runner overhead; anything in the 7–8B class wants four to five, and that is the line that doesn't fit next to everything else. The embedding model that powers memory recall needs its share, and then you still need real headroom for context, audio, and the browser engine that drives the touchscreen. There is no slack — every resident megabyte is a decision.

The rule that saved us

Treat RAM like a spreadsheet, not a vibe. Every model that stays warm is a line item with a number next to it. The one time we let two things "temporarily" coexist without doing the math, the out-of-memory killer did the math for us.

The eviction trap

The sneakiest failure wasn't a crash — it was the model runner helpfully evicting the language model to make room for an embedding batch, then reloading it on the next question. Everything still "worked." It just meant the assistant sometimes took an agonizing pause while gigabytes reloaded from storage. The fix was unglamorous: set model residency in exactly one place instead of letting each code path decide for itself, give memory-recall its own bounded slice, and warm the model at boot.

Update, September 2026: we later chose to let an idle box release the model after 30 minutes rather than hold two thirds of its RAM forever. The honest cost: the first answer after a long idle spell is slower while it reloads, then it's back to its usual pace. More in the FAQ.

What a small model is honestly good for

A 4B-class model on a Pi — we ship Gemma 3 4B — is not a frontier model, and pretending otherwise would sink the product. What it is good for: conversation, questions about your own stuff (backed by local memory and retrieval), summarizing, drafting, being a companion with real context about your world. The device always shows which brain answered and how it's running. And for the tasks where you genuinely want a big model, there's an optional bring-your-own-key mode — clearly labeled, off by default, honest that those prompts leave the box.

The trade is: a beat slower and a notch less brilliant, but completely, verifiably yours. It turns out a lot of people want exactly that trade.

Things that bit us (so they don't bite you)

  • Request-time context settings on some local runners are silently ignored — bake the context length into the model configuration itself, then verify it.
  • Audio devices renumber across reboots and kernel updates. Resolve sound hardware by what it is, never by card number.
  • A model runner can die in ways that leave its API answering health checks while every actual request hangs. Health-check the work, not the port.
  • Burn-in matters. Every build runs an automated burn-in of at least 24 hours — memory, thermals, restarts — before we trust it. Boring, and it catches things nothing else does.

How fast is AI on a Raspberry Pi 5?

The published proof page reports about 2–3 tokens per second for its documented software build. The August 2026 NVMe prototype measured 3.3–3.8 tokens per second. Those are different configurations, not interchangeable promises for the microSD build.

A short response can take tens of seconds. Prompt length, output length, the selected model and other work affect latency; a cold model adds loading time. For useful comparisons, record the hardware, storage, model, software version and whether the model was already loaded.

Real hardware and checks you can repeat

21 seconds of the real prototype, unedited. This clip shows the hardware and interface; it is not a network-isolation test. See the offline verification procedure for that check.

To check disconnected operation, select Private mode, disconnect external integrations and remove the device's internet connection while keeping it powered. Ask a new question, recall a saved memory and inspect a cited local document. Use the verification guide for the connected packet-capture check. Read unedited transcripts to see actual answers, including mistakes.

DIY Raspberry Pi AI assistant vs. PHNTM One

A DIY build lets you choose every component and software project. PHNTM One combines the hardware and assistant software in a device built and supported by PHNTM. Both approaches still face the memory and speed limits of a Raspberry Pi 5.

On a small screen, scroll the table sideways to compare both options.

Building your own assistant or reserving PHNTM One
What to compareDIY Raspberry Pi buildPHNTM One
PartsChoose and source the Pi, screen or other interface, storage, power, cooling and audio hardware.Raspberry Pi 5 with 8 GB RAM, 10.1-inch touchscreen, 256 GB microSD, active cooling, power supply, speaker, USB mic and mini keyboard. See what is included.
Setup workAssemble the hardware, install the OS and models, configure audio and connect your chosen services.Assembly and software integration are part of the product. Complete the owner setup when your unit arrives.
Assistant softwareYour chosen projects determine which chat, memory, document and voice features are available and how they work together.A touchscreen interface with local chat, memory, document search and tap-to-talk voice. Explore the included features.
MaintenanceYou select, test and maintain the OS, models and software integrations.Integrated recovery and owner-approved software updates. See how device care works and the help center.
SupportHardware vendors and the maintainers or communities of your chosen software, under their respective terms.Direct contact with the builder and a 1-year limited hardware warranty under the published terms.
CostDepends on the parts you already own and the configuration you choose. Our published BOM records $409.95 in parts and packaging as paid in August 2026; that excludes your setup time and is not a current DIY quote.$799 when ordering opens, plus applicable sales tax. No required subscription for local use. Optional cloud services use your own provider account and may add charges.

Choose DIY if you want to build and maintain your own system. Choose PHNTM One if you want the integrated device and support from its builder. The hardware does not become faster simply because it is sold as an appliance; use the performance figures above to set expectations.

PHNTM One · $799 when ordering opens

Reservations are open and free: no card, no deposit and nothing charged today. Batch 1 ships after final on-device release testing. The interactive demo simulates the interface so you can explore before deciding.

Why bother?

Because people should be able to choose where their conversations are processed and keep using their assistant when an internet service is unavailable. A dedicated local device makes that choice tangible. The trade is smaller models and slower answers on affordable hardware, with the files and operation available for inspection.

Technical references

PHNTM One is a private AI appliance being built by hand — see the specs or explore the demo or reserve free. Questions? Email me — I answer everything myself.

Ready when you are

Own your AI. Keep your words.

Reserve yours free today. Every unit is built by hand and burned in for at least 24 hours before it ships. $799 once, with no required subscription for local use.

free reservation, nothing charged · batch 1 ships after final on-device release testing · free US shipping · 30-day returns

$799 once · no subscription
Reserve