Molecule

Say what should exist. Molecule assembles a company to make it.

Inspiration

Most AI shopping tools help you find something that already exists.

We wanted to see what happens when the thing you want doesn't exist yet.

Creating a custom product usually means finding suppliers, comparing conflicting information, checking capacity, collecting quotes, coordinating production, and starting over when something changes.

We wanted that entire process to begin with one request.

What if describing a product was enough to assemble the temporary company needed to make it?

What it does

Molecule turns a spoken, typed, or multimodal request into a production network across independent merchants.

A user can ask:

“200 premium black onboarding kits by next Friday under CAD 7,000. No leather. Embroidered hoodies, engraved bottles, vegan snacks, and individual packaging.”

From that request, Molecule understands the product, processes messy supplier data, creates persistent merchant agents, collects grounded quotes and capacity, builds a dependency graph, mathematically validates the complete production plan, executes it across Shopify stores, and continuously watches the network for changes.

The user can also simply say:

“No polyester.”

Molecule creates a new version of the request without losing the budget, deadline, components, assets, or previous constraints.

If a supplier then goes offline, Molecule does not restart. It invalidates only the affected part of the graph, finds a replacement, re-solves the plan, and reconciles only the Shopify work that actually changed.

How we built it

System architecture

┌─────────────────────────────────────────────────────────────────────┐
│                         Molecule Interface                          │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐  ┌─────────┐ │
│  │Web Command   │  │macOS Dock    │  │Voice / Text  │  │Files &  │ │
│  │Center        │  │Electron      │  │OpenAI RT     │  │Images   │ │
│  └──────────────┘  └──────────────┘  └──────────────┘  └─────────┘ │
└─────────────────────────────────────────────────────────────────────┘
                                  │
                                  ▼
┌─────────────────────────────────────────────────────────────────────┐
│                     Intent Compilation                              │
│  ┌─────────────────┐  ┌─────────────────┐  ┌─────────────────────┐ │
│  │OpenAI Responses │  │Structured       │  │Intent Versioning &  │ │
│  │API              │  │Product Intent   │  │Semantic Validation  │ │
│  └─────────────────┘  └─────────────────┘  └─────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
                                  │
                 ┌────────────────┴────────────────┐
                 ▼                                 ▼
┌────────────────────────────────┐   ┌────────────────────────────────┐
│     Messy Data Agent Layer     │   │      Backboard Agent Layer     │
│                                │   │                                │
│ ┌──────────┐ ┌───────────────┐ │   │ ┌──────────┐ ┌───────────────┐ │
│ │Ingestion │ │Claim          │ │   │ │Merchant  │ │Documents +    │ │
│ │& Extract │ │Resolution     │ │   │ │Twins     │ │RAG + Memory   │ │
│ └──────────┘ └───────────────┘ │   │ └──────────┘ └───────────────┘ │
│ ┌──────────┐ ┌───────────────┐ │   │ ┌──────────┐ ┌───────────────┐ │
│ │Conflict &│ │Provenance /   │ │   │ │Typed     │ │Model Routing  │ │
│ │Quarantine│ │Uncertainty    │ │   │ │Tools     │ │+ Agent Council│ │
│ └──────────┘ └───────────────┘ │   │ └──────────┘ └───────────────┘ │
└────────────────────────────────┘   └────────────────────────────────┘
                 │                                 │
                 └────────────────┬────────────────┘
                                  ▼
┌─────────────────────────────────────────────────────────────────────┐
│                       Production Planning                           │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐  ┌────────┐ │
│  │Merchant      │  │OR-Tools      │  │NetworkX      │  │Tiger   │ │
│  │Quotes        │  │CP-SAT Solver │  │Dependency DAG│  │p95 Risk│ │
│  └──────────────┘  └──────────────┘  └──────────────┘  └────────┘ │
│                                                                     │
│           Budget • Deadline • Capacity • Materials • Risk           │
│                         VALID  /  UNSAT                              │
└─────────────────────────────────────────────────────────────────────┘
                                  │
                           User Approval
                                  │
                                  ▼
┌─────────────────────────────────────────────────────────────────────┐
│                         Shopify Execution                           │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────┐  ┌────────┐ │
│  │Composite     │  │Customer      │  │Supplier      │  │Action  │ │
│  │Product       │  │Draft Order   │  │Work Orders   │  │Journal │ │
│  └──────────────┘  └──────────────┘  └──────────────┘  └────────┘ │
│                                                                     │
│               Idempotent • Reconciled • Replay-safe                 │
└─────────────────────────────────────────────────────────────────────┘
                                  │
                                  ▼
┌─────────────────────────────────────────────────────────────────────┐
│                    Tiger Data / Durable State                       │
│ ┌────────────┐ ┌──────────────┐ ┌──────────────┐ ┌───────────────┐ │
│ │PostgreSQL  │ │TimescaleDB   │ │Continuous    │ │pgvector +     │ │
│ │State       │ │Hypertables   │ │Aggregates    │ │Trigram Search │ │
│ └────────────┘ └──────────────┘ └──────────────┘ └───────────────┘ │
│                                                                     │
│       Events • Metrics • Provenance • SSE Replay • Crash Recovery   │
└─────────────────────────────────────────────────────────────────────┘
                                  │
                                  ▼
                    ┌────────────────────────┐
                    │   Something changes   │
                    │ inventory • supplier  │
                    │ capacity • lead time  │
                    └───────────┬────────────┘
                                │
                                ▼
                    ┌────────────────────────┐
                    │   Invalidate only the  │
                    │   affected graph nodes │
                    └───────────┬────────────┘
                                │
                                └───────────────▶ Re-plan → Reconcile

The most important architectural decision was separating AI interpretation from deterministic certification.

Models can understand, extract, retrieve, and recommend.

Only our constraint solver can mark a production plan VALID.

How we use OpenAI

OpenAI powers the main customer interaction and intent-compilation layer of Molecule.

We use the Responses API with Structured Outputs to turn natural language, images, PDFs, and other context into a typed production intent containing components, transformations, quantities, materials, dietary restrictions, budget, deadline, assets, and dependencies.

We do not trust the result just because it is valid JSON. We validate it again and run semantic graph checks for cycles, duplicate nodes, missing producers, impossible transformations, and broken dependencies before planning begins.

Corrections are versioned, so saying “No polyester” updates the current production intent instead of rebuilding it from scratch. Quotes or solver responses generated from an older version cannot overwrite the new one.

For the macOS dock, we use gpt-realtime-2.1 over WebRTC for realtime interaction and gpt-4o-mini-transcribe for transcription. OpenAI credentials stay server-side, with only ephemeral authorization reaching the client.

We also intentionally restrict what the model can do. Voice can update the request and interact with Molecule, but it cannot approve commerce or directly call Shopify.

How Codex helped us

I also used Codex heavily while building Molecule, especially for debugging, validating edge cases, and speeding up iteration across the OpenAI, solver, and integration layers.

One concrete example was a real solver regression around zero capacity. Our original Python logic could treat a confirmed capacity of 0 like an unknown value, allowing an unavailable supplier to remain eligible.

We fixed the boundary condition and expanded our adversarial tests around zero capacity, cyclic graphs, invalid dependencies, currency mismatches, merchant timeouts, stale state, and duplicate execution attempts.

That directly improved the correctness of the final system.

And I'll let this image speak for itself:

OpenAI Codex usage while building Molecule

How we approached messy data for Rox

We built a messy-data agent pipeline specifically around unstructured information, incomplete datasets, conflicting sources, noisy operational data, uncertainty, and meaningful downstream actions.

It works across supplier emails, PDFs, spreadsheets, messages, call transcripts, support records, OCR output, inventory snapshots, APIs, Shopify events, scraped pages, and legacy exports.

For large-scale extraction we use GPT-5.4 Mini with Structured Outputs, while text-embedding-3-small provides 1536-dimensional embeddings for entity matching. Harder ambiguous cases can escalate to a higher-reasoning model instead of being forced into a supplier.

Every extracted fact has to include the actual evidence supporting it. If the evidence cannot be found in the source, the claim is rejected.

Entity matching progressively moves through:

exact matching → containment → trigram similarity → HNSW vector search → model adjudication

Ambiguous cases are allowed to stay unresolved.

After extraction, deterministic logic handles units, currencies, dates, lead-time periods, and capacity. Conflicting claims are compared using source authority, recency, confidence, and corroboration.

The important part is that the agents do something with the result.

A confirmed capacity shortage can trigger a re-plan. A stale or conflicted fact can trigger a supplier clarification request. Resolved values can flow back into the commerce layer after approval.

On our 1,152-artifact synthetic regression benchmark with 691 ground-truth facts, the agent achieved 86.0% extraction precision and 94.4% recall, compared with 53.6% and 37.6% from our regex baseline.

It also achieved 100% prompt-injection defense and 0% claimed-evidence-missing on that evaluation.

On a separate set of 334 capacity extractions, our earlier approach generated 69 incorrect claims. Our strict uncertainty policy reduced that to 0, preferring review over making something up.

How Backboard powers our Merchant Twins

Every supplier in Molecule becomes a persistent Merchant Twin running on Backboard.

We use Backboard across the full merchant-agent lifecycle: assistants, per-order threads, document indexing, RAG, persistent memory, tool calling, structured responses, model discovery, and multi-agent reasoning.

Each merchant gets its own identity and context instead of sharing one marketplace-wide prompt.

Supplier policies, capability documents, and merchant-specific information are indexed into that Merchant Twin. Cross-order memory lets the merchant retain information such as:

“Never auto-accept rush embroidery above 40 units while machine #2 is down.”

Backboard tool calls ground each merchant against live Molecule data for capability policies, supplier facts, current capacity, inventory, and quote calculations.

Those tools are read-only while quoting, so an agent cannot silently change inventory or operational state.

We also built model routing on top of Backboard's live catalog.

Our router discovered 16,325 models in the live catalog and groups work into four lanes:

Fast Operations, Bulk Extraction, High Reasoning, and Vision.

For harder decisions, Molecule can launch a three-agent council with separate operations, risk, and contract perspectives before returning a structured recommendation.

During live integration, we also found that combining document RAG with strict structured output required a separate final structured-response turn. We adapted the Merchant Twin architecture around that instead of giving up either RAG or typed output.

We live-verified assistant creation, document indexing, memory creation and retrieval, and identity replay without creating duplicate merchant assistants.

Deterministic production planning

Merchant agents provide grounded information and quotes, but no agent is allowed to certify feasibility.

Molecule uses OR-Tools CP-SAT to solve the production plan and NetworkX to validate the dependency graph.

hoodie ──▶ embroidery ──▶ embroidered hoodie ─┐
                                              │
bottle ──▶ engraving ───▶ engraved bottle ────┼──▶ assembly
                                              │
vegan snacks ─────────────────────────────────┘
                                                    │
                                                    ▼
                                             individual packaging
                                                    │
                                                    ▼
                                                delivery

The solver evaluates inventory, quantity, supplier throughput, setup fees, unit cost, minimum orders, material restrictions, product attributes, dependencies, historical p95 lead time, supplier availability, budget, and deadline.

The result is either:

VALID or UNSAT.

There is no LLM-generated “this probably works.”

How Shopify turns the plan into real commerce

Shopify is not a checkout API added at the end of Molecule.

It is the commerce execution layer that turns the production graph into actual merchant operations.

We use the Shopify Admin GraphQL API, version 2026-07, across 8 Shopify development stores.

For an approved plan, Molecule uses productSet to create or update the composite finished product and attach planning, risk, and provenance information through Shopify metafields.

We use draftOrderCreate for both the final customer order and individual supplier work orders. Each production node becomes a concrete obligation inside the correct supplier's Shopify store.

This means the graph is represented by actual Shopify commerce objects instead of only existing inside our UI.

Execution is also idempotent.

Every Shopify mutation is journaled before the API call and given a deterministic identity. If a request times out, the process crashes, or an execution is replayed, Molecule reconciles against existing Shopify state instead of blindly creating another object.

We live-tested this on a development store. After creating a composite product and draft orders, we restarted the process and replayed the same execution.

Zero additional Shopify mutations were created.

Live testing also exposed Shopify's 40-character draft-order tag limit, so we changed our execution identifiers to match the provider's real constraints instead of relying only on mocks.

Shopify also closes the recovery loop.

An HMAC-verified inventory update webhook can become new operational evidence. If inventory invalidates a supplier, Molecule can re-run CP-SAT, preserve unaffected supplier jobs, supersede only the invalid work, and create the replacement.

The loop becomes:

Shopify change → operational evidence → re-plan → reconciled Shopify execution

Across our configured stores, we live-read 20,815 products and 22,238 variants.

How Tiger Data powers state, analytics, and recovery

Tiger Data is the durable data and analytics backbone of Molecule.

We use Tiger Cloud with PostgreSQL, TimescaleDB, Timescale Toolkit, and pgvector so relational marketplace state and high-frequency operational streams can live together and still be queried with normal SQL.

Our Tiger environment contains approximately 5.0 million order lines and 1.2 million fulfillment samples.

We use Timescale hypertables for operational events, fulfillment history, market metrics, model-call telemetry, and large commerce datasets.

We then use Continuous Aggregates for merchant health, capability capacity, daily commerce analytics, agent cost, and supplier lead-time distributions.

For supplier risk, Timescale Toolkit's percentile_agg calculates p50, p95, and p99 lead times.

That p95 value is not just displayed in a dashboard. It feeds directly back into CP-SAT when Molecule chooses suppliers.

Our measured 12-year commerce query improved from:

8.2 seconds → 42 milliseconds

using a Continuous Aggregate.

That is a 196x speedup with identical results.

Supplier p95 analysis across 227 suppliers also improved from 3.1 seconds to 1.7 seconds.

Timescale compression reduced our order-line dataset from approximately:

1.29 GB → 352 MB

for 6.7x overall compression and up to 9.9x on compressed chunks.

Tiger also supports the messy-data pipeline through pgvector, HNSW vector indexes, trigram similarity, and standard relational SQL for entity resolution.

Finally, Tiger stores Molecule's durable event history.

Each committed event receives an ordered cursor. The frontend uses those cursors for SSE replay, so after a browser disconnect or orchestrator restart, the UI resumes from the last persisted state instead of losing the company being assembled.

Tiger ends up acting as our relational database, time-series engine, analytics layer, vector layer, and durable event backbone.

Challenges we ran into

The hardest problem was deciding what AI should not be allowed to decide.

We ended up giving each part of the system one clear responsibility:

OpenAI interprets. Our messy-data agents extract reality. Backboard represents merchants. Tiger preserves and measures state. CP-SAT proves feasibility. Shopify executes.

We also learned that external execution completely changes how retries need to work.

Once real commerce objects exist, a timeout cannot safely mean “try again.” That forced us to design around idempotency, reconciliation, provenance, plan versions, and durable state from the beginning.

Accomplishments we're proud of

Molecule works as one connected loop:

request → evidence → merchant agents → quotes → production graph → solver → approval → Shopify execution → state change → recovery

Some of the results we measured while building it:

20,815 Shopify products / 22,238 variants read live · 16,325 Backboard models discovered · 5.0M order lines / 1.2M fulfillment samples in Tiger · 196x Continuous Aggregate speedup · 6.7x–9.9x compression · 86.0% precision / 94.4% recall on our messy-data benchmark · 100% prompt-injection defense · 69 → 0 incorrect capacity claims · zero duplicate Shopify mutations after fresh-process replay.

The part we like most is still the failure case.

When a supplier disappears after the company has already been assembled, Molecule keeps going.

What we learned

Reliable AI systems are not created by giving one model more control.

They become stronger when every layer has clear boundaries, uncertainty is represented honestly, and real-world actions are recoverable.

What's next

We want merchants to connect their actual Shopify stores, operational documents, inventory systems, production tools, and fulfillment history directly to their Merchant Twin.

From there, Molecule could dynamically assemble production networks from real businesses instead of a seeded marketplace.

The long-term idea stays the same:

Say what should exist. Molecule assembles the company to make it.

Built With

  • backboard
  • codex
  • electron
  • fastapi
  • gpt-4o-mini-transcribe
  • gpt-realtime
  • networkx
  • next.js
  • node.js
  • openai-realtime-api
  • openai-responses-api
  • or-tools-cp-sat
  • postgresql
  • pydantic
  • python
  • react
  • react-flow
  • shopify-admin-graphql-api
  • sse
  • tiger-data
  • timescale-toolkit
  • timescaledb
  • typescript
  • webrtc
Share this project:

Updates

Submission history