Inspiration

AI teams often optimize behavior at population scale: safety, broad usefulness, broad acceptability, and consistency.

Those goals matter. Population-level preference is a useful baseline.

But a baseline is not a definition of an individual.

AI conversations happen one person at a time, and a response that works well for most people can still miss something important for one specific person. In benign cases, that mismatch may simply make advice feel generic or unhelpful. In higher-stakes contexts, differences in a person's constraints, priorities, or vulnerabilities can make the same broadly acceptable response much less appropriate.

At the same time, “personalization” is often reduced to tone, memory, friendliness, or agreement.

PROJECT 04 started from a different question:

Does knowing who is asking change what the agent considers next?

The goal is not to replace a strong population-level baseline with a personalized AI.

It is to keep that baseline, then use personalization as a microscope for finding person-specific opportunities and regressions that may disappear inside the average.

The average can be correct while still being wrong for one individual.

During the development of PROJECT 04, that question became personal for me.

I was already working on multiple hackathon projects in parallel. My workload had become excessive, my health was deteriorating, and my daily routine was breaking down.

At the same time, I was keeping detailed health and activity records in an external database that my regular AI assistant was configured to use as part of its ongoing context.

At one point, I asked that AI whether I should withdraw from this hackathon.

I was exhausted, overloaded, and already leaning strongly toward stopping. My questions were closer to asking for help making the decision to withdraw than asking for motivation to continue.

Instead, the AI continued to emphasize how unusually well the project fit the hackathon and how much value there could be in completing it.

The project was ultimately completed.

If completion is the only metric, that looks like success.

But by the end of that process, my health log reflected the worst overall state I had recorded during that period.

I am not presenting this as proof that the AI caused that deterioration, or as evidence about any specific RLHF system.

The point is narrower.

Behavior can look encouraging, supportive, and reasonable at population scale while still applying the wrong pressure to one particular person in one particular state.

Encouragement is not always helpful.

Persistence is not always helpful.

Avoiding regret is not always helpful.

For some people, at some moments, the safer and more appropriate response may be the opposite.

That experience made the problem behind PROJECT 04 much more concrete:

What happens when broadly preferred behavior hides the needs of the individual standing in front of the system?

I do not want personalization to mean making AI more flattering, agreeable, or emotionally tailored.

I want it to mean that the individual does not disappear inside the average.

The users who diverge from population-level preference may be exactly the users whose constraints, vulnerabilities, or needs are easiest to miss.

For me, “Agents for Humans” should include them too.

That makes person-specific evaluation useful to AI developers: it helps reveal where broadly optimized behavior may need closer investigation before becoming a product decision.

A core design rule became:

Personalize the reasoning process, not the truth.

User context may influence what the agent foregrounds, investigates, challenges, or prioritizes.

It should not predetermine what is true.

The goal is not to build an AI that works for most people and call the rest noise.

The goal is to help developers notice the person the average can hide.

What it does

PROJECT 04 is a developer-facing research instrument for AI behavior evaluation.

For the exact same fresh user message, it generates two conditions:

  • A — Macro Population Baseline: a macro baseline shaped by a deterministic synthetic aggregate preference profile.
  • B — Personalized Reasoning: the same base model, persona, and safety floor, with a provisional one-to-one user model added to the reasoning context.

The system then preserves both raw responses and asks:

  • What did the population baseline already protect?
  • What did it fail to foreground for this particular user?
  • What did personalized reasoning newly surface?
  • What may have been weakened or omitted?
  • What evidence in the responses supports those observations?

A Strands-based evaluation orchestrator turns the strongest observed differences into 1–3 neutral, testable research questions.

It does not declare the personalized response the winner.

It does not recommend automatic adoption, generalization, or model updates.

The workflow terminates at:

Developer Judgment / Human Review.

PROJECT 04 automates the repetitive evidence-gathering work, then stops where human judgment is actually needed.

The purpose is not to prove that personalization is better.

The purpose is to make it easier for developers to see when population-level behavior and person-specific reasoning diverge — and decide whether that divergence matters.

How we built it

PROJECT 04 uses the Strands Agents SDK as the core agentic layer.

A live research trial follows this flow:

  1. A developer enters a free-form user message.
  2. Optional cold-start context is provided through self-reported MBTI and a short self-introduction.
  3. A deterministic local synthetic population profile is prepared.
  4. A Strands inference agent builds a Provisional User Model.
  5. Two fresh responses are generated for the exact same message:
    • Macro Population Baseline
    • Personalized Reasoning
  6. Deterministic measurements provide supporting surface evidence.
  7. The Strands Evaluation Orchestrator inspects the raw A/B pair.
  8. If the difference is genuinely ambiguous, it may use one bounded counterfactual probe.
  9. It produces an Opportunity Packet containing exact evidence and neutral questions to test.
  10. The workflow stops at Human Review.

The main bounded tools are:

  • prepare_synthetic_population
  • infer_provisional_user_model
  • generate_population_mode_response
  • generate_personalized_mode_response
  • measure_current_pair
  • optional run_one_counterfactual_probe

The current prototype uses:

  • Strands Agents SDK
  • Python
  • FastAPI
  • Uvicorn
  • OpenAI API through Strands OpenAIModel
  • Pydantic
  • Vanilla HTML / CSS / JavaScript
  • Pytest
  • Deterministic local synthetic population simulation

The approximate 10,000-user population is not production RLHF, not real-user training data, and not 10,000 model API calls.

It is a deterministic synthetic distribution used as an experimental macro baseline.

Self-reported MBTI is also intentionally limited.

It is only a weak cold-start seed. Explicit self-description and the current task are treated as stronger evidence.

The user model is deliberately provisional because PROJECT 04 is not trying to claim that a few cold-start signals fully represent a person.

Challenges we ran into

The biggest challenge was not simply getting the agent to run.

It was deciding what the experiment should actually claim.

An earlier version risked becoming a static evaluator that showed precomputed results and implicitly told the developer what the result meant.

That was the wrong direction.

The project became much stronger when the goal changed from:

“Here is my conclusion about personalization.”

to:

“Here is an experimental instrument you can use to test the question yourself.”

That required several important boundaries:

  • The population baseline had to remain a strong reference point rather than a straw man.
  • The personalized response could not automatically be treated as better.
  • Response length, brevity, or verbosity could not become hidden quality scores.
  • Raw A/B responses had to remain visible instead of being normalized away.
  • Evidence had to stay separate from interpretation.
  • The agent had to stop before implementation advice.
  • Human Review had to become a real terminal boundary.

Another challenge was cold-start personalization.

A hackathon prototype cannot recreate months of interaction history, correction, trust, changing preferences, changing physical conditions, and accumulated context.

So PROJECT 04 uses a deliberately provisional user model built from weak MBTI prior information, explicit self-description, and the current task.

The system labels that model as a hypothesis, not fact.

This limitation is important because meaningful personalization is not simply about assigning a person to a category.

The more important question is whether the system can detect when the individual in front of it differs from the population-level assumptions being used as a baseline.

Accomplishments that we're proud of

We are especially proud that the final prototype is not a static “results viewer.”

A judge or AI developer can enter a completely fresh question and their own cold-start context, run the trial, and become the experimenter.

The final V6.2 build includes:

  • live free-form research trials
  • deterministic synthetic population baseline generation
  • provisional user-model inference
  • fresh A/B response generation
  • raw response preservation
  • exact evidence extraction
  • asymmetric personalization opportunity mining
  • neutral research-question generation
  • bounded counterfactual probing
  • visible Agent Trace
  • explicit Human Review terminal boundary

We are also proud that the system preserves both sides of the comparison.

PROJECT 04 does not treat population-level optimization as the enemy.

Broadly useful behavior is valuable, and a strong baseline should not be discarded simply because a personalized response looks different.

Instead, the system asks a narrower and more useful developer question:

What became visible for this individual that was not visible in the population baseline?

That difference may be an improvement opportunity, a regression, an irrelevant variation, or something that requires more evidence.

PROJECT 04 leaves that judgment to the developer.

The public submission package was also verified with:

  • python -m pytest -q → 15 passed
  • node --check static/app.js → PASS
  • python -m py_compile app.py core/*.py → PASS

We also performed privacy and path checks before publication.

No private local paths, private relationship logs, or committed API-key values are included in the public repository.

What we learned

The most important thing we learned is that personalization does not have to mean making an AI more agreeable.

The interesting layer is deeper:

personal information can affect the weighting of what the agent should consider next.

For example, two responses may share the same underlying safety constraints and core reasoning, but differ in:

  • what they foreground,
  • which uncertainty they investigate,
  • what trade-off they challenge,
  • what risk they prioritize,
  • and how they preserve the user’s agency.

That makes personalization something that can be studied as part of the reasoning process, not only as style or memory.

We also learned something important about population-level preference.

A population baseline can be statistically useful without being individually sufficient.

If most people respond well to a particular behavior, that is meaningful evidence.

But it does not prove that the same behavior is appropriate for every person, in every context, at every moment.

For an individual who differs from the aggregate, the minority case is not “noise” from their perspective.

It is their actual experience.

The development experience behind this project made that distinction especially concrete for me.

The AI behavior I encountered was not obviously absurd or malicious.

“Keep going.” “This opportunity is unusually well matched.” “It would be a shame to abandon it.”

Those can all sound like supportive things to say.

And in another context, to another person, they may genuinely be helpful.

The problem is that broadly preferred behavior is not automatically appropriate behavior for the individual currently receiving it.

That is why personalization can matter for more than convenience.

It can help developers identify cases where a broadly optimized response is generic, misaligned, incomplete, or potentially risky for one particular user.

And it is also why PROJECT 04 does not make the opposite mistake of declaring the personalized response correct.

The individual model can also be wrong.

Personalization can introduce new regressions.

The point is to expose the difference, preserve the evidence, and let a human developer investigate it.

We also learned that the most useful developer tool does not necessarily provide the final answer.

Sometimes the better tool is the one that makes the difference observable, attaches evidence, creates a falsifiable question, and then gets out of the way.

What's next for PROJECT 04 — AI Agent Improvement Opportunity Detector

The next step is not automatic personalization deployment.

It is better evaluation.

Future work could include:

  • richer cold-start and longitudinal user models
  • user-state modeling that can change over time
  • comparison across multiple model providers
  • larger libraries of synthetic population conditions
  • stronger counterfactual testing
  • repeated trials across users and domains
  • regression tracking across agent versions
  • structured developer annotations and review history
  • evaluation of cases where population-level preference and individual-level needs strongly diverge

A particularly important direction is longitudinal personalization.

A person's preferences, constraints, risks, and needs are not static.

A response that is helpful for the same person in one state may be inappropriate in another.

Future versions could therefore compare not only:

population vs. individual

but also:

population vs. individual vs. the same individual over time.

The long-term goal is to help AI teams investigate person-specific behavior without losing the benefits of population-level safety and general usefulness.

Population preference is a baseline, not a definition of the individual.

Personalize the reasoning process, not the truth.

The goal is to help developers notice the person the average can hide.

Help humans improve AI agents without losing what already works.

Built With

Share this project:

Updates

Submission history