Inspiration

Every AI tells you your idea is great. That's useless.I already use AI as an honest thinking partner — interviewing myself, demanding the hard read instead of the agreeable one, hunting for the pattern I didn't name. Most people don't have that habit and wouldn't know how to prompt their way into it. Whetstone hands them the habit with the hard parts pre-built.

The core engineering problem is anti-sycophancy — making GPT-5.6 challenge assumptions, cite the user's own words back at them, and concede when the user is genuinely right, instead of defaulting to praise.

What it does

Discovery interviews you — one question at a time, opening with a mood check in your own words that shapes everything after. It digs past polished answers for concrete moments and follows wherever your energy spikes.

The Mirror is the payoff: a letter that names where your energy spiked and where it went flat, quotes you to yourself, and surfaces one pattern you never named. Generic praise is banned in the system prompt. Every Mirror saves with a mood stamp, building an Archive you can compare across moods over time.

The Workshop: bring a raw idea. The first response opens with a challenge — never a compliment — and every piece of pushback carries a visible citation to your own Mirror. You can't get defensive against yourself. When your answer genuinely resolves the concern, it says so and moves on. Four exchanges, an honest close.

How we built it

Day one of this repo contains zero code — a locked build spec, a machine-readable guardrails file (AGENTS.md: scope limits, a one-table schema ceiling, banned features, error-handling rules), and the system prompts as versioned files. Codex wrote effectively every line of application code against that target, with no local runtime — every change verified through the deploy pipeline. The Vercel log is the honest record: 25+ deployments on day one, eight consecutive failures before first green.

Because anti-sycophancy is the product, it's enforced in code, not hoped for in prompts: a planted-flaw eval harness runs six test ideas — each hiding one bad assumption — through the production Workshop route and asserts no praise openers and a Mirror citation on every first response. A failing eval blocks any prompt or model change. And where prompts weren't enough, behavior moved into code: the mood→interview pivot, interview caps, and Workshop caps are deterministic.

Challenges we ran into

Authentication went through three iterations in one evening — magic-link PKCE broke across browsers, token-hash verification next, then OTP codes, then replaced wholesale with Clerk once real-world testing kept exposing redirect fragility. A TypeScript 7 / Next.js incompatibility produced a genuinely misleading error that burned three deploys before Codex diagnosed the actual cause and pinned the version. And early live sessions caught GPT-5.6 turning a "slept terribly" mood answer into five turns of medical-adjacent probing — fixed structurally, not with prompt pleading.

Accomplishments that we're proud of

The eval suite is the one I'd show first. Anti-sycophancy is usually a vibe; here it's a test. Six planted-flaw ideas run through the production Workshop route, asserting no praise-openers and a Mirror citation on every first response — PASS: 6/6, and a failing eval blocks any prompt or model change. The product's core claim is falsifiable, and a judge can run it.

The Mirror passed the only test that matters — on its builder. Four real sessions produced four different, accurate letters with zero praise adjectives. One independently diagnosed my documented failure mode — chasing small patches instead of the root fix — from seven turns of conversation. It told me something true I didn't want to hear, which is the entire product promise, working on the person hardest to surprise.

A real novel cracked open in four exchanges. The first full Workshop session took a stalled book premise and — by challenging assumptions and citing my own Mirror — surfaced the organizing connection I'd been missing for weeks: the family fracture and the magical fracture are the same event. It never suggested a plot point. It asked the questions that made me find it. Ship criterion met with 48 hours to spare: a stranger can register with an email code and hold a saved, mood-stamped Mirror in under six minutes — live, on a real URL, with visible error states and no silent failures.

And the process itself: day one of the repo was zero code — spec, guardrails, prompts. Codex built the rest through 25+ deployments, three auth redesigns, and a misleading TypeScript 7 error, while I ran a full week of accelerator and investor meetings in parallel. The spec held; the scope never crept.

What we learned

Whetstone v1 took me weeks on a no-code platform — no tests, no evals, no real auth. This rebuild took days with Codex: real auth, persistence, defensive JSON validation, visible error states, and a passing eval suite born the same day as the prompt it tests. The delta isn't just speed — it's that spec-driven agent development produced a more rigorous product than hand-assembled tooling did.

What's next for Whetstone

A live idea canvas that fills as a Workshop runs (returning from v1), Mirror-over-time comparison views, and optional gentle re-engagement — the current Archive-only retention is a deliberate choice.

Built With

Share this project:

Updates

Submission history