scroll
thriving
Production Solo-built

Fogo

Track your spending, stick to your budget, build toward financial independence. An AI companion that reads your actual financial state. Not a budget app listing transactions. Not a chatbot wrapping them.

An AI personal finance app, solo-built. This case study is the April 2026 snapshot of the architecture. The June to August 2026 sprint that followed is summarized on the portfolio, and the live product is at askfogo.com.

A five-tier classification cascade routes transactions through rules and a merchant cache before the API. Every LLM call receives a SQL-generated financial state vector and logs it to Langfuse. A two-layer eval harness scores every response on data grounding, tone, actionability, insight, and conversation. Currently distilling the production system into an open-source model.

FastAPI React PostgreSQL Anthropic API Langfuse Plaid Railway Vercel
01 Plaid sync Bank-connected, live transactions. Not a demo database. 02 Full picture Income and spending together, tracked against real budgets. 03 Agentic budgets Chat creates and adjusts budgets. 9-tool registry, real commits. 04 Review loop Monthly AI review. Tone shifts with your financial situation.
01 · Intelligence Layer

SQL builds the context. A state machine sets the tone. The LLM gets both.

Nine rules evaluated before any API call. The model gets pre-computed context and tone guidance, not raw data to interpret.

Facts before the call. Tone before the call. SQL assembles the structured state object, and a deterministic state machine maps it to an emotional state with tone guidance injected into the system prompt. The app responds differently when you're burning through budget versus thriving because the system told it to, not because the LLM decided to. Both financial_context and emotional_state are first-class fields on every Langfuse trace.

State machine · 9 conditions evaluated
first match wins · no LLM involvement
1 pace_ratio > 1.40 → pacing MATCH
2 sr < -0.10 → stressed SKIP
3 adherence < 0.40 → concerned SKIP
4 anomaly > 0.60 → concerned SKIP
5–9 … NOT REACHED
Financial context
assembled from SQL · no LLM calls
pace_ratio1.41spending ahead of budget pace
sr-0.08savings rate this month
adherence0.62categories on track / total
anomaly0.30unusual charges flagged
days_elapsed0.86day 26 of 30
9 rules evaluated in priority order
EMOTIONAL STATE
pacing
INTENSITY
0.71
TONE GUIDANCE · injected into system prompt
"Lead with what they can still control.
No alarm without a path forward."
02 · Transaction Classification

Every incoming bank transaction runs a four-stage cascade. Most never reach the LLM.

The system teaches itself to stop calling the API. After three consistent classifications for a merchant, a rule is promoted and that transaction type is handled for free, forever.

Four tiers, cheapest first. Most merchants never reach the LLM. A merchant rules table hits first (zero cost). Then a cache of prior LLM results. Then Claude Haiku. Then human triage. After three high-confidence results for a merchant, a rule is auto-promoted and that merchant exits the cascade permanently.

Classification and display are two separate problems. The LLM assigns one of ~70 canonical labels (food_delivery, salary_paycheck) and that label is permanent. Budget groups and display subcategories are computed on-the-fly and never persisted, so changing how a transaction is grouped never reclassifies it. Invalid labels fall back to uncategorized at zero confidence rather than persisting bad data.

Classification cascade

Plaid tx TRADER JOE'S #527 −$54.18 ↓
1
Merchant rulesExact match lookup
HIT
2
Merchant cachePreviously seen names
HIT
3
Claude HaikuLLM with cost routing
4
Triage queueHuman review fallback
UBER* EATS food_delivery 99%
SQ *BLUE BOTTLE coffee_cafe 94%
AMZN MKTP US online_shopping 71%
ACH TRANSFER 4821 — triage !
Canonical taxonomy · 101 permanent labels
TYPE
Spending
Income
Transfer
Savings
Debt
BUDGET GROUP dynamic
Dining & Delivery
Transport
Healthcare
Shopping
user-configurable
CANONICAL LABEL stored
food_delivery
coffee_cafe
restaurant_dining
fast_food
permanent · ~70 targets
03 · Ask Fogo: Agentic Tool Use

Talk with it, query it, have it recategorize. One conversation.

It queries your actual transactions, moves charges between categories, and commits the change in one conversational turn.

Nine tools. Real queries, real writes, real commits. Ask what happened to a budget category and it runs a query against your actual transactions. Ask it to move charges and it calls the right tool and commits the change. It understands your full financial state (pace, savings rate, adherence) and links across categories when answering.

The registry is the stable surface; the model is a replaceable component. The agentic work extends beyond chat: opaque merchant disambiguation (resolving Amazon and Apple charges by reasoning over order history) is a multi-step pipeline built on the same tool registry. The conversation is one interface to it.

Ask Fogo
Where did my dining budget go this month?
query_transactions · dining · march 2026

You've spent $412 on dining. That's $162 over your $250 budget. The breakdown:

Uber Eats
$198
Chipotle
$89
DoorDash
$72
Other
$53
Move the Uber Eats charges to food delivery
update_transaction_category · Uber Eats · 8 charges

Done. 8 Uber Eats charges reclassified to food_delivery. Dining is now $214. Back under your $250 budget.

04 · Evaluation Framework

Two layers, five dimensions, and a rubric built from watching real responses fail.

Dimension one requires no LLM. It's a programmatic check that the model didn't hallucinate a number or leak an internal label into user-facing text.

Two layers: the first requires no LLM at all. Layer 1 is deterministic: did it call the right tool, stay within latency thresholds, avoid leaking internal labels like transfer_internal into user-facing text? These checks block quality scoring on failure. Layer 2 is quality: five dimensions scored by a Sonnet judge with chain-of-thought, calibrated at 1, 3, and 5.

Five dimensions, each targeting a distinct failure mode. Data grounding catches hallucinated amounts. Tone calibration catches factually correct but emotionally off responses. Actionability catches vague advice. Insight depth catches responses that restate the dashboard. Conversation skill catches responses that read like a report. Each is weighted differently per surface: tone calibration matters most for the review, insight depth for analytical chat. The eval is a living artifact: when the system changes, the rubric is what gets updated.

Deeper dive: how I calibrated the judge →
Layer 1 · programmatic checks no LLM
numbers match SQL output PASS
no internal labels in output PASS
tool called with correct params PASS
latency under threshold PASS
Eval result
review_narrate · pacing · march 2026
01 data_grounding PASS programmatic
02 tone_calibration 4.1 / 5 sonnet judge
03 actionability 3.8 / 5 sonnet judge
04 insight_depth 4.6 / 5 sonnet judge
05 conversation_skill 4.2 / 5 sonnet judge
emotional_state pacing · intensity 0.71
05 · Training Data Flywheel

Every production call logs a complete training tuple.

Every production LLM call automatically emits a complete training tuple. No separate data collection step.

Every prod call is already a training example. Every LLM call writes a Langfuse trace with three labeled fields: the normalized financial state vector, the deterministic emotional state, and the model output. No separate collection step. Thumbs up/down on chat messages produces DPO pairs linked directly to trace IDs.

Every agentic pipeline that ships is also a labeling pipeline. Opaque merchant resolution produces ground-truth examples (charge matched to actual item category) that feed directly into the training set. Each tuple is (state_vector, emotional_state, output, preference) with no ambiguity about the conditioning variable. Target: distillation onto a smaller open-source model via DPO/KTO.

Langfuse trace
ambient · dashboard · march 2026
financial_state { pace_ratio: 1.41, sr: -0.08,
adherence: 0.62, anomaly: 0.3 }
emotional_state pacing · intensity 0.71

model claude-haiku-4-5
tokens_in / out 480 / 94
cost_usd $0.00031
latency_ms 847

feedback thumbs up · linked to trace_id
async def call_llm(prompt, surface, user_id, *, financial_context=None):
    # surface → rate limits + routing + Langfuse trace name
    # financial_context → attached to every trace as first-class field
Roadmap

Live, building, and what comes after.

Fogo is live in production on Railway and Vercel.

● Live in Production
Railway · Vercel · running today
  • +
    Full web app: Dashboard, Transactions, Budget, Analytics, Ask FogoSpending · income · savings rate lenses. Real data throughout, no synthetic fallbacks.
  • +
    Two-phase monthly review: reflect, then planHolds you accountable to last month's goals. Sets next month's budget in the same session.
  • +
    Classification cascade: rules → cache → Haiku, auto-promotionMost merchants exit the cascade before reaching the LLM.
  • +
    Nine function-calling tools, LLM gateway with cost tracking + LangfuseRead-only to Haiku, writes to Sonnet.
  • +
    Deterministic emotional state machine: 9 states, logged on every traceSame context always maps to same state. Reproducible.
  • +
    Five-dimension eval: dim 1 programmatic, dims 2–5 Sonnet judge
  • +
    Training data on every LLM call: state vector + output + preference pairs
◆ Building Now
Active development
  • →
    User onboarding flowFirst-run experience: account connection, transaction history pull, AI-suggested budget initialization.
  • →
    RLHF training loopDataset accumulating in Langfuse. DPO/KTO pipeline is the immediate next step.
  • →
    Opaque merchant resolutionAmazon/Apple/Venmo disambiguation. Dual purpose: better product + ground-truth training data.
  • →
    Open-source model distillationFine-tune a smaller model via DPO/KTO: cheaper to serve, better on this domain.