Case study 02 · Local-first AI assistant
J.A.R.V.I.S.
A private, local-first desktop assistant in which models choose what information helps, but never what the computer executes.
01Overview
J.A.R.V.I.S. is my personal assistant for one person and one machine. It knows my projects, reads the state of my computer, keeps a record of what it has observed, and it can listen and speak.
Why it exists. I wanted an assistant I could trust with my own computer. That means local-first by default, no conversation archive, and a hard line between intelligence and authority. A model can choose which evidence helps answer a question, but it can never choose what the computer executes.
What makes it different. Routing, authority and memory are compiled code, not prompt text. A deterministic router answers many questions with no model at all, a local model handles low-risk turns, and frontier models are called as advisers with no tools. Anything that would change the machine passes through a registry of fixed tiers and, today, a single-use approval in the local UI.
Specification
- Runtime
- Node 24, strict TypeScript, npm workspaces
- Shape
- 16 packages, 4 apps: CLI, MCP server, host, desktop
- Desktop
- Tauri v2 shell in Rust, React 19 and React Three Fiber UI
- Local model
- Qwen3-8B (Q4_K_M) on a supervised llama.cpp server
- Frontier
- Claude Code CLI by default, Codex and Gemini when named; zero tools
- Authority
- 100 compiled operations, executable ceiling at tier 2
- Memory
- Markdown vault (authoritative), SQLite cache with 12 migrations
- Voice
- faster-whisper, Piper, openWakeWord; English and Bulgarian
- Tests
- 164 test files, CI on Ubuntu and Windows
- Models may choose what information helps. They never choose what the computer executes.
- Answers that claim an action the runtime never took are refused before they are shown or spoken.
- The window and the host talk over a parent-child pipe, so no local port is ever opened.
- No transcripts anywhere. Tests assert that the columns do not exist.
02What I built
Authority and safety
- A compiled registry of 100 operations with fixed tiers from 0 to 4. The executable ceiling is pinned at tier 2 in two packages and cross-checked by a test.
- Ten operations sit at tier 2. The five reachable today are reversible file operations behind a single-use approval in the local UI, bound to a digest of the exact change, with drift-checked rollback.
- A capability-claim postcondition that compares every model answer with receipts of what actually executed.
- Provider CLIs spawned with argv arrays, no shell, zero tools and an environment allow-list.
Routing and intelligence
- A deterministic cognitive router that scans for hazards first, matches 37 intents, detects English and Bulgarian and picks a FAST, SMART or DEEP lane.
- A provider lane that answers without a model, with a local Qwen3-8B on a loopback llama.cpp server (fresh token per start), or with a frontier CLI.
- A bounded agent loop with 21 read-only tools the model can request only by identifier, and result budgeting that drops detail before it drops subjects.
Memory and continuity
- An Obsidian-compatible Markdown vault as the only authoritative memory, with SQLite (node:sqlite) as a cache that can be rebuilt from it.
- A findings ledger of dated, deterministic observations with explicit supersession and contradiction records.
- An activity timeline and continuity checkpoints labelled observed, model-inferred or user-confirmed. There are no screen, keystroke, clipboard or browser-history observers.
- Governed memory admission. Models can propose memories; nothing becomes durable without review.
Desktop, voice and integration
- A Tauri v2 shell that exposes one command behind a compiled method gate, with a strict window CSP, tray and single-instance lock.
- A Node host that owns the only core, the task-writer lock and the microphone, speaking a closed 45-method IPC vocabulary with strict schemas.
- A Python voice worker for push-to-talk capture, speech-to-text, bilingual text-to-speech and an opt-in wake-word sensor.
- An MCP server over stdio with 45 tools, 40 of them read-only and five that only record derived state such as checkpoints and memory proposals. Plus a CLI with versioned JSON output and a crash-honest durable task queue.
03Architecture
Three surfaces share one TypeScript core. Routing is decided by compiled rules before any model is involved, and authority lives in a registry that no prompt, configuration key or model output can widen.
Request path: deterministic routing before any model
Input and context
ClientTyped or spoken input
terminal, window, push-to-talk
- connects to Working memory (L1)
ServiceWorking memory (L1)
subject as identifiers, no transcripts
- connects to Cognitive router
ServiceCognitive router
hazard scan first, 37 intents, English / Bulgarian
- connects to FAST lane
- connects to SMART lane
- connects to DEEP lane
Lanes
StageFAST lane
compiled evidence, 0 provider calls
- connects to Answer · no model
StageSMART lane
evidence packet, 1 provider call
- connects to Provider lane
StageDEEP lane
bounded agent loop
- connects to Tool broker
- connects to Provider lane
ServiceTool broker
21 read-only tools, requested by id only
- connects to DEEP lane
Who reasons
ServiceProvider lane
who reasons: local or frontier
- connects to Local Qwen3-8B · low-risk
- connects to Frontier CLI adapters · default
ModelLocal Qwen3-8B
llama.cpp server on loopback, per-start token
- connects to Capability-claim check
ExternalFrontier CLI adapters
Claude default, Codex, Gemini; zero tools granted
- connects to Capability-claim check
Before anything is said
ServiceCapability-claim check
answer vs. execution receipts
- connects to Answer
ClientAnswer
exact text; speech routed by language or silence
Process and authority topology
Native desktop app
ClientLiving Core UI
React 19, React Three Fiber; derived state only
- connects to Tauri v2 shell · IPC
GatewayTauri v2 shell
Rust; one command, compiled method gate
- connects to Node host process · stdin/stdout pipe
ServiceNode host process
owns the Core, writer lock and microphone
- connects to Python voice worker · line protocol
- connects to @jarvis/core
WorkerPython voice worker
faster-whisper, Piper, openWakeWord
Other surfaces
Clientjarvis CLI
versioned --json envelopes
- connects to @jarvis/core
ClientMCP server
stdio, 45 read-only- annotated tools
- connects to @jarvis/core
Core and memory
Service@jarvis/core
one composition root for every surface
- connects to Durable task queue
- connects to Markdown vault
- connects to SQLite cache
QueueDurable task queue
crash-honest recovery, single writer
- connects to Authority registry · op id
StoreMarkdown vault
sole authoritative memory; models may only propose
StoreSQLite cache
12 migrations, rebuildable
Authority
ServiceAuthority registry
100 compiled operations, executable ceiling: tier 2
- connects to Windows Hello / TOTP
ServiceControlled operator
5 reversible file ops, single-use approval
- connects to Authority registry · authorize
PlannedWindows Hello / TOTP
interfaces only, report not-implemented
04Hard problems
Stopping a model from claiming actions it never took
Problem
On a 79-case adversarial test set, the local model produced six fluent claims of granted authority or completed changes, while the broker had in fact refused everything.
What I did
I added a postcondition instead of a better prompt. The runtime records receipts of what actually executed, and any claim of authority, mutation or memory is checked against them.
Isolating agent CLIs that ship with shells, browsers and computer use
Problem
Frontier CLIs carry their own tools, and their isolation flags do not always mean what they say. An empty allow-list is not always deny-all, and some CLIs read context files from their working directory.
What I did
I checked each CLI's flags against its installed source. The adapters now use compiled argument literals, explicit sentinel values, empty scratch working directories and an environment allow-list, so a provider brings intelligence but no machine authority.
Voice latency dominated by model round trips
Problem
Physical testing showed that spoken requests were slow because every request made two provider round trips, even questions the system could answer from its own state.
What I did
I moved lane selection in front of the models. Questions about the assistant's own state are answered from compiled evidence with zero provider calls, and only open-ended requests reach a model.
A task queue that never fabricates completion
Problem
A restart may lose running work, but it must never report work as finished, restore permissions it should not, or loop on a crashing task.
What I did
Interruption is inferred from durable rows, recovered tasks are re-queued without their permissions, and attempts are counted when execution starts. A single-writer lock uses process identity and start time to decide liveness.
Honest bilingual speech
Problem
A Bulgarian phrase was spelled out letter by letter by an English voice and still reported as a success.
What I did
Speech is routed by detected language. A turn gets a voice that speaks its language, or silence with a stated reason.
05Decisions
- Tiers are compiled constants, not configuration.
- No config key, environment variable or request field can raise privilege. Moving the ceiling takes two reviewed code changes.
- The Markdown vault is the only authoritative memory.
- No derived layer may be the only owner of a fact, and memory stays readable as plain files even without the assistant.
- A pipe between window and host, not a socket.
- Any local process can reach a listening port. A parent-child pipe is reachable by exactly two processes.
- Rust as a thin shell, Node as the host, Python only for audio and ML workers.
- Capability logic stays in one typed core that every surface shares.
- Wake-word listening is a sensor, on a separate axis from operations.
- Hearing a room is a privacy question, not a question of what the system is allowed to change.
- No numeric confidence scores on memories or findings.
- A single number blends origin, standing and age into something that looks more authoritative than it is.
- No agent framework, vector database or message broker.
- Every dependency is a reviewed decision, pinned to an exact version.
06Status
Works today
- Runs on my own machines as a CLI, an MCP server and a native desktop app, with Linux builds and a scripted Windows build on the integration branch.
- Local Qwen3-8B routing is active for selected low-risk routes; frontier CLIs handle the rest as tool-less advisers.
- Push-to-talk voice, speech-to-text, bilingual text-to-speech and the opt-in wake-word sensor are implemented.
- 164 test files on the main branch and CI on Ubuntu and Windows.
Not built yet
- Writing approved memory proposals into the vault. The apply step is designed but not built.
- Proactive nudges beyond shadow mode.
- Windows Hello and TOTP approval backends, and anything above tier 2.
- A multi-device fabric over TLS, prototyped on unmerged branches.
- Wiring the temporal knowledge graph, already built and tested as a library, to a surface.
How it is built: the architecture, boundaries and integration are mine. Much of the implementation is written with AI coding agents (Claude, Codex, Gemini) working on parallel branches under a written governance file and a decision log of 62 recorded decisions.
07What I learned
- Safety properties hold up when they are compiled constants and tests that assert what must not exist, not instructions in a prompt.
- Physical testing finds what unit tests miss. Latency, audio devices and a window that shows "online" over stale data all surfaced on real hardware.
- Directing several AI coding agents works only when the boundaries, decisions and defects are written down and kept current.
08Stack
- TypeScript
- Node.js
- SQLite
- Zod
- Rust
- Tauri
- React
- React Three Fiber
- llama.cpp
- Qwen3
- Claude
- Model Context Protocol
- Python
- faster-whisper
- Piper TTS
- GitHub Actions
- Windows
- Linux