-
-
Groundwork
-
The pipeline
-
The report that will give you
-
Get a resilience score, your tasks sorted into *under pressure* and *holding*
-
A phased plan where **every milestone cites the task id and figure that justifies it
-
The nearest occupation with genuinely higher ground
-
List and summarize every tasks
-
Choose your job title
-
Enter your skill
-
Enter what matters in your work
-
Enter the kind of work pull at you
-
Enter the time you can devote to it
-
Enter your own opinion
Inspiration
A friend told me she'd been informed for two years that her job was doomed. When I asked who told her, the answer was: headlines, and ChatGPT.
So I asked ChatGPT the same question. It gave back a confident, well-written paragraph with a percentage in it. I asked where the percentage came from. There was no answer — because it hadn't come from anywhere. The number and the explanation had been produced by the same generation step, at the same instant.
That's a fine failure mode for trivia. It is not fine for something someone will resign over.
The second realisation was that the question itself is framed wrong. Nobody replaces a data scientist. They replace writing the weekly cleaning script. An occupation is forty-one distinguishable tasks, and exposure across them is wildly uneven — so the useful answer was never a verdict. It's a list.
What it does
Groundwork breaks a job into its real O*NET tasks and scores each one against two independent published sources: the Anthropic Economic Index, which records AI usage that has actually been observed, and Eloundou et al. (2023), a peer-reviewed theoretical exposure coefficient.
A task only scores high when both sources agree. Where they disagree, the score lands between them and neither is allowed to win.
You get a resilience score, your tasks sorted into under pressure and holding, the nearest occupation with genuinely higher ground, and a phased plan where every milestone cites the task id and figure that justifies it.
The defining property is what it refuses to do: no language model touches the number. Four stages, exactly one calls a model, and it's the last one — writing sentences about arithmetic that already happened.
How I built it
React 18 + Vite front end, FastAPI + pandas back end, Featherless.ai for the two model calls, deployed on Vercel and Render. No database, no accounts — the questionnaire lives in browser memory and is never stored.
The data layer joins O*NET 29.0, the Economic Index, and the Eloundou coefficients: 922 occupations, 18,747 scored task rows.
Scoring, per task: composite = (economic_index_weight + β) / 2 when both measures exist, β alone when usage was never observed. Per occupation, weighted by O*NET's incumbent importance × frequency ratings.
The agent runs a hand-rolled JSON action loop over five computed tools rather than OpenAI-style function calling — per-model support for that parameter is unverifiable across Featherless's 20,000+ model catalogue and fails as an opaque 400 when missing.
Three things are enforced in code, not requested in a prompt:
- every number the model writes is discarded and replaced with the computed one
- any milestone citing a task id that doesn't exist is dropped; if none survive, the whole plan falls back to the computed version
- a destination may never score lower than where the person already is
Challenges I ran into
A bug that only appeared once the product worked. Every non-ASCII character from the model arrived mangled — 3.5â¯hrs, atârisk. SSE responses carry no charset, so requests silently falls back to ISO-8859-1 and decodes UTF-8 a byte at a time. It stayed invisible for the entire build because the computed fallback path is pure ASCII, and it would have hit every user the moment the model path went live. One line: resp.encoding = "utf-8".
The report contradicted itself. The page printed "25 of your 31 tasks" (computed) directly above the model's own "1 at-risk task out of 27". The prompt had asked it to state the finding in figures — so it recalled numbers it had seen, and got them wrong. On a product whose entire claim is that its numbers aren't written, that was the worst possible thing to leave on screen. The fix was to forbid the summary from restating counts at all.
The plan recommended a worse job. Asked to plan a move, the agent picked a destination scoring 39 against the 44 the person already had, because its interest fit was 80.4% versus 79.4%. It traded five points of the thing being measured for one point of a tie-breaker — which the prompt had invited by saying to prefer better interest fit without bounding the cost.
Accomplishments that I'm proud of
The claim is checkable in thirty seconds. grep the three scoring modules for any model or network call and you get nothing. Compute a real score and print whether the model client was even imported — False. Grep the built JS bundle for the API key — 0. All three are in the README, because "trust me" was the exact failure I set out to fix.
The score survives the model failing. It resolves in ~6ms while the agent takes 15–25 seconds, and nothing clears it when plan generation fails. "Your score survived. The plan didn't." is a literal description of the state machine, not a slogan.
I kept none and unknown apart. One means usage was measured and was minimal — evidence. The other means nothing was measured — absence of evidence. Conflating them would have scored 81% of the corpus as fully resilient. Most of the honesty in this product is downstream of that one distinction.
I publish our my worst number. Mean observed-usage coverage is 17.7%, median 11.8%, and 232 occupations have none at all. Every report prints its own coverage rather than hiding behind an average.
What I learned
Prompts are requests. Code makes guarantees. Every instruction we gave the model was eventually violated in some run — invented figures, uncited milestones, a destination that made things worse. The only rules that held were the ones enforced after the model spoke.
A grounded product holds you to its own standard. Our landing page carried six testimonials from six people who don't exist — on a page promising "no made-up numbers." I'd written the thing I built the product to criticise, and didn't notice until someone read it back. Deleting them mattered more than any feature: a reader who works out that Maya Ellsworth isn't real has no reason to believe the 18,747 either.
Motion is where honesty leaks. The score gauge rewound to zero before requesting its first animation frame — so in a print, a background tab, or with reduced motion enabled, it displayed a confident 0. Not a missing number. A wrong one.
The most valuable output was the constraint. Forcing every milestone to name a task id didn't just make plans checkable — it made them better. A model that has to cite something stops writing "learn AI" and starts writing "move one weekly cleaning job behind a reviewed model diff."
What's next for Groundwork
Close the coverage gap honestly. The Economic Index updates quarterly. Each release raises observed-usage coverage above today's 17.7% mean — and because coverage is printed on every report, that improvement is visible to users rather than silent.
Semantic retrieval over the task corpus. Designed, deliberately not built. It would let the agent search all 18,747 task statements by meaning — answering "which occupations involve work like this" instead of navigating occupation by occupation. Shelved because an embedding index is a second source of truth to keep honest, and the grounded path works without one.
Cross-occupation comparison in the UI. The /compare endpoint already exists and returns shared tasks and the gap between any two occupations. Nothing renders it yet.
Re-run reminders. A score is a measurement of the present, so it goes stale. Six months is the right interval — the Index updates quarterly, and a stale score is worse than no score.
What I won't add: accounts, stored answers, or a forecast. The privacy claim is only credible because there is no database to leak, and the moment this predicts job loss rather than measuring observed usage, it becomes the thing it was built against.

Log in or sign up for Devpost to join the conversation.