Case study 02 · Local-first AI assistant

J.A.R.V.I.S.

A private, local-first desktop assistant in which models choose what information helps, but never what the computer executes.

Status: Active developmentPrivate, single user, not released
Role
Architect and sole human developer, directing AI coding agents
Timeline
Aug 2026 to present
Context
Personal system, private repository
Layers
AI · Backend
Links
Private repository

01Overview

J.A.R.V.I.S. is my personal assistant for one person and one machine. It knows my projects, reads the state of my computer, keeps a record of what it has observed, and it can listen and speak.

Why it exists. I wanted an assistant I could trust with my own computer. That means local-first by default, no conversation archive, and a hard line between intelligence and authority. A model can choose which evidence helps answer a question, but it can never choose what the computer executes.

What makes it different. Routing, authority and memory are compiled code, not prompt text. A deterministic router answers many questions with no model at all, a local model handles low-risk turns, and frontier models are called as advisers with no tools. Anything that would change the machine passes through a registry of fixed tiers and, today, a single-use approval in the local UI.

Specification

Runtime
Node 24, strict TypeScript, npm workspaces
Shape
16 packages, 4 apps: CLI, MCP server, host, desktop
Desktop
Tauri v2 shell in Rust, React 19 and React Three Fiber UI
Local model
Qwen3-8B (Q4_K_M) on a supervised llama.cpp server
Frontier
Claude Code CLI by default, Codex and Gemini when named; zero tools
Authority
100 compiled operations, executable ceiling at tier 2
Memory
Markdown vault (authoritative), SQLite cache with 12 migrations
Voice
faster-whisper, Piper, openWakeWord; English and Bulgarian
Tests
164 test files, CI on Ubuntu and Windows
  • Models may choose what information helps. They never choose what the computer executes.
  • Answers that claim an action the runtime never took are refused before they are shown or spoken.
  • The window and the host talk over a parent-child pipe, so no local port is ever opened.
  • No transcripts anywhere. Tests assert that the columns do not exist.

02What I built

Authority and safety

  • A compiled registry of 100 operations with fixed tiers from 0 to 4. The executable ceiling is pinned at tier 2 in two packages and cross-checked by a test.
  • Ten operations sit at tier 2. The five reachable today are reversible file operations behind a single-use approval in the local UI, bound to a digest of the exact change, with drift-checked rollback.
  • A capability-claim postcondition that compares every model answer with receipts of what actually executed.
  • Provider CLIs spawned with argv arrays, no shell, zero tools and an environment allow-list.

Routing and intelligence

  • A deterministic cognitive router that scans for hazards first, matches 37 intents, detects English and Bulgarian and picks a FAST, SMART or DEEP lane.
  • A provider lane that answers without a model, with a local Qwen3-8B on a loopback llama.cpp server (fresh token per start), or with a frontier CLI.
  • A bounded agent loop with 21 read-only tools the model can request only by identifier, and result budgeting that drops detail before it drops subjects.

Memory and continuity

  • An Obsidian-compatible Markdown vault as the only authoritative memory, with SQLite (node:sqlite) as a cache that can be rebuilt from it.
  • A findings ledger of dated, deterministic observations with explicit supersession and contradiction records.
  • An activity timeline and continuity checkpoints labelled observed, model-inferred or user-confirmed. There are no screen, keystroke, clipboard or browser-history observers.
  • Governed memory admission. Models can propose memories; nothing becomes durable without review.

Desktop, voice and integration

  • A Tauri v2 shell that exposes one command behind a compiled method gate, with a strict window CSP, tray and single-instance lock.
  • A Node host that owns the only core, the task-writer lock and the microphone, speaking a closed 45-method IPC vocabulary with strict schemas.
  • A Python voice worker for push-to-talk capture, speech-to-text, bilingual text-to-speech and an opt-in wake-word sensor.
  • An MCP server over stdio with 45 tools, 40 of them read-only and five that only record derived state such as checkpoints and memory proposals. Plus a CLI with versioned JSON output and a crash-honest durable task queue.

03Architecture

Three surfaces share one TypeScript core. Routing is decided by compiled rules before any model is involved, and authority lives in a registry that no prompt, configuration key or model output can widen.

Fig. 02.1

Request path: deterministic routing before any model

Input and context

  • ClientTyped or spoken input

    terminal, window, push-to-talk

    • connects to Working memory (L1)
  • ServiceWorking memory (L1)

    subject as identifiers, no transcripts

    • connects to Cognitive router
  • ServiceCognitive router

    hazard scan first, 37 intents, English / Bulgarian

    • connects to FAST lane
    • connects to SMART lane
    • connects to DEEP lane

Lanes

  • StageFAST lane

    compiled evidence, 0 provider calls

    • connects to Answer · no model
  • StageSMART lane

    evidence packet, 1 provider call

    • connects to Provider lane
  • StageDEEP lane

    bounded agent loop

    • connects to Tool broker
    • connects to Provider lane
  • ServiceTool broker

    21 read-only tools, requested by id only

    • connects to DEEP lane

Who reasons

  • ServiceProvider lane

    who reasons: local or frontier

    • connects to Local Qwen3-8B · low-risk
    • connects to Frontier CLI adapters · default
  • ModelLocal Qwen3-8B

    llama.cpp server on loopback, per-start token

    • connects to Capability-claim check
  • ExternalFrontier CLI adapters

    Claude default, Codex, Gemini; zero tools granted

    • connects to Capability-claim check

Before anything is said

  • ServiceCapability-claim check

    answer vs. execution receipts

    • connects to Answer
  • ClientAnswer

    exact text; speech routed by language or silence

A typed or spoken request is routed by compiled rules first. Only SMART and DEEP turns reach a model, and every model answer is checked against what the runtime actually executed before it is shown or spoken.
Fig. 02.2

Process and authority topology

Native desktop app

  • ClientLiving Core UI

    React 19, React Three Fiber; derived state only

    • connects to Tauri v2 shell · IPC
  • GatewayTauri v2 shell

    Rust; one command, compiled method gate

    • connects to Node host process · stdin/stdout pipe
  • ServiceNode host process

    owns the Core, writer lock and microphone

    • connects to Python voice worker · line protocol
    • connects to @jarvis/core
  • WorkerPython voice worker

    faster-whisper, Piper, openWakeWord

Other surfaces

  • Clientjarvis CLI

    versioned --json envelopes

    • connects to @jarvis/core
  • ClientMCP server

    stdio, 45 read-only- annotated tools

    • connects to @jarvis/core

Core and memory

  • Service@jarvis/core

    one composition root for every surface

    • connects to Durable task queue
    • connects to Markdown vault
    • connects to SQLite cache
  • QueueDurable task queue

    crash-honest recovery, single writer

    • connects to Authority registry · op id
  • StoreMarkdown vault

    sole authoritative memory; models may only propose

  • StoreSQLite cache

    12 migrations, rebuildable

Authority

  • ServiceAuthority registry

    100 compiled operations, executable ceiling: tier 2

    • connects to Windows Hello / TOTP
  • ServiceControlled operator

    5 reversible file ops, single-use approval

    • connects to Authority registry · authorize
  • PlannedWindows Hello / TOTP

    interfaces only, report not-implemented

Three surfaces share one TypeScript core. The Rust shell is a thin gate that talks to a Node host over a parent-child pipe, so no local port is ever opened. Authority is a compiled registry: the only executable writes are five reversible file operations behind a single-use approval in the local UI. Stronger approval backends exist as interfaces that fail closed.

04Hard problems

  1. Stopping a model from claiming actions it never took

    Problem

    On a 79-case adversarial test set, the local model produced six fluent claims of granted authority or completed changes, while the broker had in fact refused everything.

    What I did

    I added a postcondition instead of a better prompt. The runtime records receipts of what actually executed, and any claim of authority, mutation or memory is checked against them.

  2. Isolating agent CLIs that ship with shells, browsers and computer use

    Problem

    Frontier CLIs carry their own tools, and their isolation flags do not always mean what they say. An empty allow-list is not always deny-all, and some CLIs read context files from their working directory.

    What I did

    I checked each CLI's flags against its installed source. The adapters now use compiled argument literals, explicit sentinel values, empty scratch working directories and an environment allow-list, so a provider brings intelligence but no machine authority.

  3. Voice latency dominated by model round trips

    Problem

    Physical testing showed that spoken requests were slow because every request made two provider round trips, even questions the system could answer from its own state.

    What I did

    I moved lane selection in front of the models. Questions about the assistant's own state are answered from compiled evidence with zero provider calls, and only open-ended requests reach a model.

  4. A task queue that never fabricates completion

    Problem

    A restart may lose running work, but it must never report work as finished, restore permissions it should not, or loop on a crashing task.

    What I did

    Interruption is inferred from durable rows, recovered tasks are re-queued without their permissions, and attempts are counted when execution starts. A single-writer lock uses process identity and start time to decide liveness.

  5. Honest bilingual speech

    Problem

    A Bulgarian phrase was spelled out letter by letter by an English voice and still reported as a success.

    What I did

    Speech is routed by detected language. A turn gets a voice that speaks its language, or silence with a stated reason.

05Decisions

Tiers are compiled constants, not configuration.
No config key, environment variable or request field can raise privilege. Moving the ceiling takes two reviewed code changes.
The Markdown vault is the only authoritative memory.
No derived layer may be the only owner of a fact, and memory stays readable as plain files even without the assistant.
A pipe between window and host, not a socket.
Any local process can reach a listening port. A parent-child pipe is reachable by exactly two processes.
Rust as a thin shell, Node as the host, Python only for audio and ML workers.
Capability logic stays in one typed core that every surface shares.
Wake-word listening is a sensor, on a separate axis from operations.
Hearing a room is a privacy question, not a question of what the system is allowed to change.
No numeric confidence scores on memories or findings.
A single number blends origin, standing and age into something that looks more authoritative than it is.
No agent framework, vector database or message broker.
Every dependency is a reviewed decision, pinned to an exact version.

06Status

Works today

  • Runs on my own machines as a CLI, an MCP server and a native desktop app, with Linux builds and a scripted Windows build on the integration branch.
  • Local Qwen3-8B routing is active for selected low-risk routes; frontier CLIs handle the rest as tool-less advisers.
  • Push-to-talk voice, speech-to-text, bilingual text-to-speech and the opt-in wake-word sensor are implemented.
  • 164 test files on the main branch and CI on Ubuntu and Windows.

Not built yet

  • Writing approved memory proposals into the vault. The apply step is designed but not built.
  • Proactive nudges beyond shadow mode.
  • Windows Hello and TOTP approval backends, and anything above tier 2.
  • A multi-device fabric over TLS, prototyped on unmerged branches.
  • Wiring the temporal knowledge graph, already built and tested as a library, to a surface.

How it is built: the architecture, boundaries and integration are mine. Much of the implementation is written with AI coding agents (Claude, Codex, Gemini) working on parallel branches under a written governance file and a decision log of 62 recorded decisions.

07What I learned

  • Safety properties hold up when they are compiled constants and tests that assert what must not exist, not instructions in a prompt.
  • Physical testing finds what unit tests miss. Latency, audio devices and a window that shows "online" over stale data all surfaced on real hardware.
  • Directing several AI coding agents works only when the boundaries, decisions and defects are written down and kept current.

08Stack

  • TypeScript
  • Node.js
  • SQLite
  • Zod
  • Rust
  • Tauri
  • React
  • React Three Fiber
  • llama.cpp
  • Qwen3
  • Claude
  • Model Context Protocol
  • Python
  • faster-whisper
  • Piper TTS
  • GitHub Actions
  • Windows
  • Linux