We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

Inspiration

Development teams increasingly build features in parallel. Two changes can pass review and testing independently, merge without a Git conflict, and still break the product because they rely on incompatible assumptions.

That gap inspired Jointly.

Our demonstration uses a checkout system where one change adds percentage coupons while another adds payment retry validation. Both changes pass their tests, and Git merges them cleanly. However, the coupon changes the meaning of Order.total, while payment retry still assumes the original accounting rule. Only a test derived from both changes’ intended behavior reveals the collision.

What it does

Jointly is an intent-aware pre-merge verification system for independently developed changes.

It:

  1. Freezes two changes and their common Git base.
  2. Creates isolated base, change-A, change-B, and combined workspaces.
  3. Runs the existing test suites independently and together.
  4. Converts each change’s intent into structured requirements.
  5. Identifies high-risk interaction surfaces.
  6. Generates a focused cross-change test.
  7. Distinguishes real behavioral collisions from invalid tests, setup errors, timeouts, and environmental failures.
  8. Evaluates a repair in the isolated combined workspace.
  9. Re-runs interaction, regression, and stability checks.
  10. Produces an auditable Merge Safety Passport.

The principle behind Jointly is:

AI proposes. Evidence decides.

A model can suggest a hypothesis, test, or repair, but Jointly does not accept model confidence as proof. Every verdict must be supported by deterministic execution and recorded evidence.

How we built it

Jointly is a TypeScript and Node.js monorepo containing:

  • a deterministic verification core;
  • four-workspace Git isolation;
  • structured intent, hypothesis, test, repair, and evidence contracts;
  • typed MCP tools for bounded repository operations;
  • a provider-independent reasoning layer;
  • a watsonx.ai transport adapter with mocked and offline contract coverage;
  • a Vite evidence dashboard;
  • Zod validation for inputs and artifacts;
  • Vitest regression and failure-classification tests;
  • a gated Merge Safety Passport.

IBM Bob assisted with parts of the planning and implementation workflow, while Jointly’s own verification core controls repository operations, executes tests, records evidence, and determines verdicts.

The website turns the investigation into a clear evidence narrative: change overview, intent map, hidden collision, verified repair, passport, and roadmap.

Challenges we faced

The hardest challenge was determining what counts as real evidence.

A generated test can fail because of a missing import, invalid syntax, broken setup, timeout, skipped execution, or environmental problem. None of those failures proves that two changes are incompatible.

We implemented structured execution records and explicit classifications so that only a meaningful assertion failure tied to the stated interaction hypothesis can confirm a collision.

Another challenge was keeping repository operations safe and reproducible. Commands are allowlisted, source branches remain untouched, and repair candidates are evaluated only inside an isolated combined workspace.

We also made a deliberate effort to distinguish demonstrated capabilities from future work. The deterministic verification core, evidence dashboard, supported local investigation workflow, and offline-tested watsonx adapter exist today. Live hosted inference, credential-separated cloud execution, multi-user authorization, and approved integration-PR publication remain roadmap milestones.

What we learned

We learned that merge safety is not only a text-merging problem. It is a problem of reconciling assumptions.

Traditional tests usually verify what each contributor anticipated independently. Jointly asks the missing question:

What test would neither contributor think to write alone?

We also learned that trustworthy AI-assisted development requires a clear boundary between probabilistic proposals and deterministic decisions. AI can help form hypotheses, but reproducible execution must determine whether those hypotheses are true.

What’s next

Our next milestone is an authorized hosted workflow that connects live watsonx.ai reasoning and GitHub repository selection to the same deterministic verification core.

Before any hosted release, Jointly will require:

  • credential-separated execution with canary-secret tests;
  • filesystem, process, and network isolation;
  • authorization between users, repositories, runs, and artifacts;
  • exact candidate and source-ref revalidation;
  • explicit human approval before publishing an integration pull request;
  • regression protection for the existing local workflow.

Longer term, Jointly can support additional languages, test frameworks, repository profiles, and interaction graphs involving more than two changes.

Jointly: Two green PRs. One hidden collision. Evidence before merge.

What's next for Jointly

Built With

Share this project:

Updates

Submission history