Vision One Million
An agent pipeline that keeps Waterloo Region's Ready 1 Million scorecard current across five domains, built at the FCI x LangChain hackathon in March 2026 and running on a weekly schedule since. This piece covers the fallback chain that keeps the numbers moving, and the failure mode that chain hides.
Context
The scorecard tracks how a region of half a million people is doing against its growth targets: homes built, transit ridership, emergency room waits, employment rate, greenhouse gas reduction. Five domains, each with its own metrics and its own target. Before the pipeline, someone updated a spreadsheet by hand.
The data exists in public. It sits with Statistics Canada, the Ontario Data Catalogue, CMHC, Grand River Transit, Ontario Health, and a climate dashboard, and each of them publishes on its own schedule in its own format. One serves a clean API. One posts a CSV. One buries the figure in an HTML table that gets restyled without warning. One ships a PDF.
That variety is the whole problem. A pipeline that assumes an API gets three metrics and stalls on the rest.
The decision that mattered
The pipeline reaches every source through a ranked chain, and search is the floor rather than the plan.
Each source declares how it should be read: an API or CSV endpoint where one exists, a scraper where the number lives in a page, a PDF extractor where it lives in a document. Those are the primary paths, ordered by how much you can trust what comes back. A parsed API response is a fact. A scraped table is a fact until the page is restyled. A figure pulled out of a PDF by a model is an inference, and it is labelled as one.
Underneath all three sits Tavily. A primary path that throws, or returns without a result, sends the fetcher to a web search for the same regional metric and takes the most recent published figure it can find. fetch_with_fallback catches both cases, the exception and the quiet empty return, because a source that answers with nothing breaks a scorecard the same way a source that answers with a stack trace does.
Every result carries a source_used field recording which tier answered: primary, openai_extraction, or tavily. The number and its provenance travel together.
The reasoning is that a scorecard that stops updating is worthless. A regional dashboard is a thing people check once a quarter, at best. If a URL changes in April and the pipeline halts, nobody finds out in April. They find out in September, when someone opens the dashboard for a meeting and the newest figure is five months old, and at that point the tool has taught them not to trust it. A figure from a web search is worse than a figure from Statistics Canada. Both are better than a blank.
Validation sits behind the chain rather than in front of it. Pydantic models type every domain's payload, and GPT-4o-mini compares each new value against its history and flags anomalous jumps for human review rather than dropping them. The pipeline is allowed to be wrong in a way a person can see. It is not allowed to be silent.
The whole thing runs from a GitHub Actions cron at midnight on Sundays, writes to SQLite, and commits the database back to the repository. Streamlit reads that file.
What it cost
Committing a binary database on a schedule is not a design anyone would defend on a whiteboard. It works here because the file is small, the write is weekly, and one commit per run gives the project a free audit trail: eighteen commits between 29 March and 26 July, one every seven days, no gaps. A hackathon build got persistence, history, and deployment out of a choice made because there was no time to provision a database.
The ReAct agent is the piece that shows the deadline. It answers natural-language questions about the scorecard through LangGraph with LangSmith tracing, and it sits beside the ingestion path rather than inside it. Ingestion is deterministic fetchers with a search fallback, and none of it routes through the agent. For a hackathon named after an agent framework, the agent ended up as the demo surface and the pipeline underneath it stayed boring. The boring part is the part still running.
Ten commit messages in the history appear twice. That is what a weekend of force-pushes looks like, and it is the honest cost of the schedule.
What I'd do differently
The chain solved the wrong half of the problem, and it took four months of green runs to notice.
The goal was that the scorecard should never stop updating, and it has not. The design never addressed a source that stops working while the numbers keep moving. If the Ontario Health page restructures in May, the scraper fails, Tavily answers, a figure lands in the database, the dashboard renders, and the weekly commit goes through. Nothing is broken from the outside. The metric has dropped from a government table to whatever a search engine surfaced, and it will stay there until somebody looks.
The information needed to catch it already exists. Every row records source_used, and the system health page counts how many metrics came from fallback. That count is a page you have to visit. For an unattended pipeline, a number that appears when you go looking and never otherwise is a number nobody reads, which is the same failure the whole project was built to fix, moved one level up.
The fix is small and I did not build it. The run knows how many sources fell through, so it can compare that against the previous run and fail the workflow when a source degrades for the first time or when the fallback count climbs. A red check on a Sunday morning is the alert. It costs a few lines in the workflow and it turns the fallback from something that hides a fault into something that reports one.
The lesson I would carry: a fallback is a way of continuing, and continuing is not the same as being fine. Anything designed to absorb a failure needs a second mechanism that makes the absorption visible, decided at the same time as the fallback rather than after four months of runs that all looked identical.