Paul S.

Cortex

Internal tooling for my own workflow: a local-first multi-agent system that ingests from several sources, classifies and enriches with LLMs through a tiered routing policy, and stores validated structured output in a vector database, encrypted at rest.

Context

A personal system for turning a stream of incoming material (RSS and news, finance data, meeting and message sources) into classified records and structured insights I can query later.

Problem

The problems are the ones any pipeline has: how to spend inference budget sensibly, how to get output that downstream code can rely on, and how to store personal data safely. A system that calls the best available model for every task and parses prose out of the response is fine to demo and unreliable once it runs unattended on a schedule.

What I did

  • Split the system across two hosts, a 24 GB VM serving models locally and a 4 GB VM running orchestration, connected over gRPC with streaming completions and an embedding endpoint.
  • Built a tiered model-routing policy: tasks declare a tier by importance, and the registry resolves tier plus domain to a specific model, with fallbacks.
  • Layered the pipeline as ingestion → classification → orchestration → LLM → storage, so that cheap deterministic work happens before anything reaches a model.
  • Required validated structured output at every LLM boundary: responses are parsed into a strict schema, normalized, and rejected instead of stored when they don't conform.
  • Encrypted payloads at rest with Fernet before they are written to the vector store, for both raw records and generated insights.
  • Exposed the result through a GraphQL API over the stored records.

Two hosts

The model process needs 24 GB of memory and is busy for seconds at a time. The orchestration service needs predictable availability, does mostly IO, and gets restarted whenever I change something. Running both in one process means every model load competes with the part that is supposed to stay up.

So they are separate hosts: a 24 GB VM running the model server and registry, and a 4 GB VM running ingestion, classification, storage and the API. They talk over gRPC, with completions streamed token by token and a separate endpoint for embeddings.

Model routing

Sending everything to the strongest available model is expensive where it is not slow. Deciding which category a record belongs to is a smaller problem than writing a daily summary across everything that arrived.

So tasks declare a tier by how much quality matters: validation and classification at the low tier, insight generation in the middle, strategic and summary work at the top. Routing resolves tier *and* domain to a concrete model, because the model that handles finance text best is not the one that classifies best.

TIER_MODEL_MAP = {
    "classification": {
        "tier2": "llama3",   # fast, high volume, low complexity
        "tier3": "mistral",  # better context handling
        "tier4": "gpt-4",    # only where the decision matters most
    },
    "finance":    {"tier2": "mistral", "tier3": "mistral", "tier4": "mistral"},
    "daily_note": {"tier2": "llama3",  "tier3": "mistral", "tier4": "gpt-4"},
    "default":    {"tier2": "llama3",  "tier3": "llama3",  "tier4": "mistral"},
}
Routing is a lookup table, so changing a model is a data change and each call site's intent stays visible as a tier.

The pipeline is tiered the same way: deterministic rules classify what they can, and the model is consulted only for what needs judgment. That also keeps the volume of model calls down on a host serving them locally.

The validation boundary

An LLM that returns prose is unusable as a component here. Once downstream code has to interpret free text, the parse can fail in ways nobody enumerated, and it fails quietly, storing a record that reads fine and carries the wrong values.

So every LLM boundary produces structured output that is parsed, validated against a schema, and normalized before anything is written. A response that does not conform is rejected instead of coerced into shape, because coercion produces records whose fields are populated and wrong.

{
  "domain": "finance",
  "priority": "high",
  "sentiment": "negative",
  "summary": "Concise generated statement.",
  "source_ids": ["raw-record-uuid"],
  "model": "mistral",
  "tier": "tier3"
}
Shape of a validated insight record. The enumerated fields are what make it queryable later.

Storage runs in the same order every time: payloads are encrypted with Fernet before they are written, for raw records and generated insights alike, and decrypted only when something reads them. This is personal data on machines I administer myself.

Result

  • A working local-first pipeline: ingest, classify, enrich through a routed model call, validate, encrypt, and store as queryable vectors.
  • Model choice is a routing-table entry, so swapping or adding a model does not touch call sites.
  • Inference is isolated on its own host, so the memory-hungry half can be restarted or re-hosted without disturbing the half that holds the data.
  • Structured, schema-validated output at every model boundary, which is what made the system dependable enough to run unattended.
  • A solo side project on two VMs I administer myself. There is no production scale behind it, no users besides me, and no uptime claim.

Tech

  • Python
  • Qdrant
  • Ollama
  • llama.cpp
  • gRPC
  • LangGraph / CrewAI
  • GraphQL
  • Fernet