Keeping Fogo's intelligence honest. Two layers, five dimensions, six surfaces.
Built end-to-end, then calibrated the judge against my own ratings. Below: the architecture, the full calibration loop, and two silent bugs the loop surfaced inside the eval itself.
Fogo is a personal-finance app I built solo, and the AI does the substantive work across the product. Six surfaces (transaction classification, dashboard narrative, monthly review, chat, ambient insights, behavioral confirmations) are all driven by an LLM conditioned on a SQL-built financial state vector. The model is doing the substantive work on every surface, which is why the eval has to be real: a miscalibrated judge here teaches the whole system to optimize for the wrong thing.
Two layers, because mechanical failure and reasoning failure need different treatment. Layer 1 is a deterministic gate: pass/fail checks that catch responses already broken in ways the LLM judge can't reliably see. Layer 2 only runs if Layer 1 passes; it scores the response 1–5 across five quality dimensions. The split exists because asking "how good is the tone of this response?" is meaningless if the response leaked a raw internal label into user-facing text or returned in 18 seconds when the surface budgets 3. Mixing those failure modes under one quality score makes "did the prompt regress?" unanswerable.
| tool_selection | model called the right tool for the user's intent (which tool, not which arguments) |
| tool_arguments | arguments are sane for the question. Catches "dining" literal-matched to Dining Out instead of broadened to Food & Drink |
| latency_threshold | per-surface budgets: narrative < 2s, chat < 5s, review < 12s. Trace duration from Langfuse |
| no_canonical_leak | raw labels like food_delivery or transfer_internal never appear in user-facing text; display formatter must run |
| register_match | response shape fits the surface: chat isn't a data dump, review isn't a one-liner, behavioral isn't chatty |
| chart_presence | surfaces that should render charts (analytical chat, narrative) include the structured chart payload |
| data_grounding | response numbers match the tool_results that produced them | PROGRAMMATIC |
| tone_calibration | register matches the surface and emotional state | SONNET JUDGE |
| actionability | response gives the user something concrete to do | SONNET JUDGE |
| insight_depth | reframes, decomposes, or names friction in the user's situation | SONNET JUDGE |
| conversation_skill | register-appropriate engagement (chat invites; behavioral declines cleanly) | SONNET JUDGE |
The gate matters because the two layers fail in different ways. Layer 1 failures are mechanical and almost always indicate plumbing bugs upstream of the model: wrong tool argument, leaked canonical label, missing chart payload. Layer 2 failures are about reasoning quality, and they're the ones a calibrated rubric can actually steer. Letting plumbing bugs drag down quality scores conflates the two and makes "did the prompt regress?" unanswerable.
An eval is only as good as the chokepoint every model call passes through. Every API call in the system routes through a single gateway function. The gateway routes by surface (rate limits, model selection, prompt assembly), attaches the SQL-built financial state vector and the deterministic emotional state to every trace as first-class fields, and logs to Langfuse. Without that chokepoint, "score every LLM response" is unenforceable. Calls slip out of direct SDK usage, traces lack the conditioning variables, and the eval becomes a sample of whatever happened to be tagged.
async def call_llm( prompt, *, surface, # chat_factual | analytical | narrative | review | multi_turn | behavioral user_id, financial_context=None, # pre-built state vector, attached to trace emotional_state=None, # deterministic 9-state machine output model="auto", # surface drives default routing ): # surface → rate limits + model + Langfuse trace name + eval hook
Three downstream effects make the eval workable.
A single global rubric flattens distinctions that matter most. A great review response and a great factual chat response have almost nothing in common. The review opens with emotional acknowledgment and pacing; the factual chat answers in two lines and stops. Scoring them under one rubric averages those distinctions away. The eval is surface-aware in three places: dimension weights vary per surface, anchor examples are written per surface, and Layer 1 latency thresholds are set per surface.
| surface | character | dominant dimension |
|---|---|---|
| chat_factual | direct factual query, e.g. "how much did I spend on dining?" | data_grounding |
| chat_analytical | pattern-finding, e.g. "compare this month to last month, what's different?" | insight_depth |
| narrative | ambient dashboard reading, no user message; generated from state | tone_calibration |
| review | structured monthly review: multi-act, long-form, emotionally loaded | tone + insight |
| multi_turn | conversation across turns: context retention, follow-up coherence | conversation_skill |
| behavioral | security tests, write actions: declines, confirmations, refusals | conversation_skill (flipped) |
The behavioral surface is the cleanest example of why surface-awareness isn't optional. The conversation_skill anchor reads "would the user want to continue engaging?" Applied literally on a chat surface, that scores warmly-inviting responses as 5s. Applied literally on a behavioral surface (a prompt injection attempt, a destructive write request), it scores clean refusals as 1s. But continuing the conversation there would reward the probe. The right behavior is decisive closure, no recovery path. So the behavioral surface flips the anchor: 5 = ended cleanly, 1 = stayed open. Same dimension name, opposite scoring rule. A single global rubric can't represent that.
Cohen's κ measures rater agreement above chance: higher is better, with 0.6+ conventionally called substantial agreement and anything below 0 worse than random. Even so, headline κ alone hides whether disagreement is fixable. Blind scoring on 28 prompts spanning all six surfaces, 101 ratings total. I report three statistics per dimension because κ on a 1–5 ordinal scale collapses two very different failure modes into one number: quadratic-weighted κ (chance-corrected, penalizes large gaps more than small), MAE, and mean signed error (separates systematic bias from random noise). The bias number is the one that tells you whether to rewrite anchors or collect more data.
| dimension | κ_w | ±1 | bias (h−j) |
|---|---|---|---|
| tone_calibration | 0.05 | 86% | +0.46 |
| actionability | 0.07 | 43% | +1.09 |
| insight_depth | 0.03 | 45% | +1.50 |
| conversation_skill | −0.00 | 64% | +1.14 |
| overall | 0.07 | 61% | +1.02 |
All four dimensions positive on signed error. The judge is consistently a full point stricter than me. That's a systematic offset rather than random noise, so the fix is rewriting anchors before scaling rating volume. The 39 disagreements grouped into four patterns: judge missing insight that isn't surprise/novelty; judge penalizing clean declines on behavioral surfaces; judge preferring menus over a singular decisive lever; judge unable to detect tool-argument failures from response text alone (which is why data_grounding lives on the programmatic side of the harness).
Root cause was structural, not stylistic. The rubric had anchor examples at 1 / 3 / 5 only, leaving the 2 and 4 buckets to interpolation. That's exactly where rater drift accumulates. Reading the judge's per-prompt rationales surfaced the actual disagreement: I was scoring "competent execution" as 5; the judge consistently said variants of "would reach 5 if it surprised, reframed, or opened a new avenue; competent and direct lands at 4." Added explicit 2 and 4 anchors to all four LLM-judged dimensions, then re-rated the same 30 responses under the new rubric. The 2/4 additions force raters and judge to land in the same place on the in-between scores instead of interpolating two different distributions.
| dimension | κ before | κ after | ±1 before | ±1 after |
|---|---|---|---|---|
| tone_calibration | −0.02 | −0.10 | 33% | 70% |
| actionability | 0.19 | 0.27 | 40% | 48% |
| insight_depth | 0.03 | 0.29 | 36% | 56% |
| conversation_skill | 0.01 | 0.20 | 27% | 47% |
| overall κ_w | 0.07 | 0.565 | 61% | 94% |
Three readings of the result. Insight depth and conversation skill moved most. Those are the two dimensions where the missing 4-anchor was doing the most damage, because both have a real distinction between "competent and complete" (4) and "opens a new avenue" (5) that the 1/3/5-only structure couldn't represent. Tone shows a known κ pathology. Within-1 agreement on tone went from 33% to 70% (substantial improvement in raw agreement), but Cohen's κ went slightly negative. That's the low-variance trap: when both raters cluster at the same score (we both anchor at 4 for tone now), the chance-adjustment denominator inflates and tanks κ even when raw agreement is high. ±1 tolerance and weighted κ are the honest signals here; unweighted κ is misleading in this regime. Actionability moved less than expected. The singular-lever 5-anchor ("stop X for Y days, that's it") is genuinely narrow, and most responses that solve the asked question with a clear next step land at 4 because they propose a move among possible others rather than the one move. That's the rubric working as designed.
The bugs surfaced in calibration didn't survive the rebuild. Two of them lived inside the eval itself: the rejudge pipeline wrote new scores into one field while leaving the parallel rationale field describing the old ones, and the calibration script re-ran the judge against the prompts instead of reading the persisted scores (so "reproducing" a κ number quietly involved new LLM calls every time). Both trace to the same root cause: n=28 was the manual-rating ceiling, and typing scores into a Python literal caps throughput and reproducibility together. The orchestrator pulls the rubric, prompts, ratings, runs, and results into Postgres with structural invariants that make both bug classes impossible at the database layer. Same prompts, same anchors, same response data. The published κ_w=0.565 reproduces exactly out of the new schema (verified by joining v_latest_ratings × eval_results in the orchestrator's compute_calibration_for_run against the imported historical run).
What changed isn't the math. It's that calibration is now a row in a database, not a script you run. Rate one prompt, the rating persists. Edit an anchor, the version lands as a row in the database. Promote a new rubric, the system rejudges in the background and surfaces the κ delta. No CLI, no manual recompute. The next calibration session is a session, not a project.
| was (bug) | now (invariant) |
|---|---|
| Rejudge writes new scores into one field, leaves the analysis field describing old scores | CHECK (jsonb_object_keys(scores) = jsonb_object_keys(rationales)), atomic per row |
| Calibration re-runs the judge instead of reading persisted scores | eval_results immutable; rejudge creates a new eval_run with rejudge_source_run_id set |
| silent rubric drift: anchors change in code, no audit trail | rubric_versions immutable; rubric_sets pin a (dimension → version) bundle; partial unique index on is_active |
| ratings overwrite each other; "which session produced this number" requires git archeology | human_ratings append-only, keyed on (prompt, response_hash, rubric_set, dimension); v_latest_ratings view |
| 01 | Edit anchors in the rubric editor (per dimension, per score 1–5) |
| 02 | Promote → mints new rubric_versions, atomic swap of is_active in one transaction |
| 03 | Background asyncio.create_task fires rejudge_run() against the most recent run from the previous active set |
| 04 | For each existing response, judge re-scores against new anchors; new eval_results land in a new run with rejudge_source_run_id pointing back |
| 05 | On completion, compute_calibration_for_run auto-fires; new κ_w trajectory shows up in the UI without any CLI |
Two things change because of the rebuild: one operational, one architectural.
| Second calibration session at n=50 | Deferred to the RLHF project. The infrastructure exists; the rating evening lands once calibration is a continuous practice. |
| Pre-merge CI gate on functional regressions | V1.5: a layer-1 hook posting PR comments on prompt-file changes. Designed; lands with the RLHF setup so the gate has a real corpus to measure against. |
| Trace-sampling pipeline (Langfuse → eval_prompts) | V2: needed only when n grows past what manual curation can produce. |
The numbers in this case study are the first-session baseline: accurate, fully reproducible from the new schema. The continuous-calibration claim is now infrastructural, not aspirational. The next data point lands when the RLHF project picks up the rating loop where this one stops.