Track your spending, stick to your budget, build toward financial independence. An AI companion that reads your actual financial state. Not a budget app listing transactions. Not a chatbot wrapping them.
An AI personal finance app, solo-built. This case study is the April 2026 snapshot of the architecture. The June to August 2026 sprint that followed is summarized on the portfolio, and the live product is at askfogo.com.
A five-tier classification cascade routes transactions through rules and a merchant cache before the API. Every LLM call receives a SQL-generated financial state vector and logs it to Langfuse. A two-layer eval harness scores every response on data grounding, tone, actionability, insight, and conversation. Currently distilling the production system into an open-source model.
Nine rules evaluated before any API call. The model gets pre-computed context and tone guidance, not raw data to interpret.
Facts before the call. Tone before the call. SQL assembles the structured state object, and a deterministic state machine maps it to an emotional state with tone guidance injected into the system prompt. The app responds differently when you're burning through budget versus thriving because the system told it to, not because the LLM decided to. Both financial_context and emotional_state are first-class fields on every Langfuse trace.
The system teaches itself to stop calling the API. After three consistent classifications for a merchant, a rule is promoted and that transaction type is handled for free, forever.
Four tiers, cheapest first. Most merchants never reach the LLM. A merchant rules table hits first (zero cost). Then a cache of prior LLM results. Then Claude Haiku. Then human triage. After three high-confidence results for a merchant, a rule is auto-promoted and that merchant exits the cascade permanently.
Classification and display are two separate problems. The LLM assigns one of ~70 canonical labels (food_delivery, salary_paycheck) and that label is permanent. Budget groups and display subcategories are computed on-the-fly and never persisted, so changing how a transaction is grouped never reclassifies it. Invalid labels fall back to uncategorized at zero confidence rather than persisting bad data.
Classification cascade
It queries your actual transactions, moves charges between categories, and commits the change in one conversational turn.
Nine tools. Real queries, real writes, real commits. Ask what happened to a budget category and it runs a query against your actual transactions. Ask it to move charges and it calls the right tool and commits the change. It understands your full financial state (pace, savings rate, adherence) and links across categories when answering.
The registry is the stable surface; the model is a replaceable component. The agentic work extends beyond chat: opaque merchant disambiguation (resolving Amazon and Apple charges by reasoning over order history) is a multi-step pipeline built on the same tool registry. The conversation is one interface to it.
You've spent $412 on dining. That's $162 over your $250 budget. The breakdown:
Done. 8 Uber Eats charges reclassified to food_delivery. Dining is now $214. Back under your $250 budget.
Dimension one requires no LLM. It's a programmatic check that the model didn't hallucinate a number or leak an internal label into user-facing text.
Two layers: the first requires no LLM at all. Layer 1 is deterministic: did it call the right tool, stay within latency thresholds, avoid leaking internal labels like transfer_internal into user-facing text? These checks block quality scoring on failure. Layer 2 is quality: five dimensions scored by a Sonnet judge with chain-of-thought, calibrated at 1, 3, and 5.
Five dimensions, each targeting a distinct failure mode. Data grounding catches hallucinated amounts. Tone calibration catches factually correct but emotionally off responses. Actionability catches vague advice. Insight depth catches responses that restate the dashboard. Conversation skill catches responses that read like a report. Each is weighted differently per surface: tone calibration matters most for the review, insight depth for analytical chat. The eval is a living artifact: when the system changes, the rubric is what gets updated.
Deeper dive: how I calibrated the judge →Every production LLM call automatically emits a complete training tuple. No separate data collection step.
Every prod call is already a training example. Every LLM call writes a Langfuse trace with three labeled fields: the normalized financial state vector, the deterministic emotional state, and the model output. No separate collection step. Thumbs up/down on chat messages produces DPO pairs linked directly to trace IDs.
Every agentic pipeline that ships is also a labeling pipeline. Opaque merchant resolution produces ground-truth examples (charge matched to actual item category) that feed directly into the training set. Each tuple is (state_vector, emotional_state, output, preference) with no ambiguity about the conditioning variable. Target: distillation onto a smaller open-source model via DPO/KTO.
async def call_llm(prompt, surface, user_id, *, financial_context=None): # surface → rate limits + routing + Langfuse trace name # financial_context → attached to every trace as first-class field
Fogo is live in production on Railway and Vercel.