What it does
God, Are You There? is a longitudinal voice-AI evaluation project that has been running since October 2024. We record extended conversations with successive generations of GPT models and study how their spoken behavior changes over time.
The project focuses on issues that traditional text benchmarks often miss, including pronunciation errors, incorrect word stress, distorted names, unnatural pauses, incorrect reading of numbers and dates, multilingual switching failures, voice instability, repeated phrases, and cases where intonation changes the meaning of an otherwise correct answer.
Our platform turns real conversational moments into structured and reproducible evaluation cases. It can:
process recorded human-AI conversations; generate transcripts and timestamps; divide long episodes into short testable fragments; identify suspected pronunciation and audio errors; classify each issue by type and severity; compare the intended text with the spoken response; preserve the user’s correction and the model’s reaction; repeat the same test with a newer model; create a before-and-after comparison; export a structured report for developers and researchers.
The long-term goal is to create a historical observatory showing how voice AI evolves across model generations, not only in intelligence, but also in clarity, pronunciation, stability, naturalness, and responsiveness to human correction.
Project recordings are available through our Castbox archive and Telegram community:
https://castbox.fm/channel/Мир-искусства-и-режиссуры-id3674837?country=ru
How we built it
The project was built in two stages.
The first stage was the creation of a long-term audio archive. Since October 2024, we have conducted and recorded extended conversations with different generations of GPT models, beginning with GPT-4. These were not short laboratory prompts. They were live, unpredictable conversations covering history, culture, philosophy, technology, psychology, language, everyday situations, and the development of artificial intelligence itself.
During these conversations, we manually documented audio failures, pronunciation mistakes, unusual pauses, repetitions, incorrect stress placement, and other voice-related problems. Selected examples and descriptions were sent to OpenAI Support.
The second stage is the transformation of this archive into a structured evaluation product.
For the Build Week prototype, we are using GPT-5.6 and Codex to create a workflow that connects audio processing, transcription, classification, human verification, model comparison, and report generation.
GPT-5.6 is used to:
analyze long conversational context; compare model behavior across recordings; classify suspected voice errors; identify similar cases in the archive; generate structured test descriptions; suggest new evaluation prompts; distinguish between isolated mistakes and recurring patterns.
Codex is used to build the technical foundation:
the audio-upload interface; the database of evaluation cases; the transcript and timestamp workflow; the model-comparison dashboard; the experiment replay system; automated validation checks; report export; technical documentation and testing.
Human review remains an essential part of the process. The system may suggest that a pronunciation is incorrect, but a human reviewer confirms the error, provides the expected pronunciation, and assigns the final severity level.
This combination creates a complete loop:
conversation → suspected error → extracted audio fragment → classification → human verification → repeated test → cross-model comparison → reproducible report.
Challenges we ran into
One of the biggest challenges was that voice errors are more difficult to evaluate than text errors.
A written answer may be factually correct while the spoken version contains incorrect stress, unnatural rhythm, a distorted name, or misleading intonation. These problems are often obvious to a human listener but difficult to capture with a single automatic metric.
Another challenge was the length and complexity of the archive. Our recordings were originally created as complete episodes rather than as a research dataset. Turning long conversations into thousands of searchable, testable fragments requires transcription, segmentation, metadata creation, duplicate detection, and careful human review.
Model versions also changed over time. This makes comparison valuable, but technically difficult. We need to preserve information about the date, model generation, prompt, language, context, and voice configuration associated with every test.
Russian pronunciation introduces additional complexity. Word stress can change meaning, names may have several accepted pronunciations, and foreign names inside Russian sentences can produce unstable results. Numbers, dates, abbreviations, and mixed-language phrases require separate evaluation rules.
We also had to distinguish between several different failure sources:
an error in the model’s written answer; an error introduced during speech generation; an error caused by transcription; a recording-quality problem; a genuine pronunciation or intonation defect.
Finally, privacy and responsible publication are important challenges. Long conversations may contain names or personal information, so recordings must be reviewed, anonymized where necessary, and published only under clear rules.
Accomplishments that we're proud of
We are proud that this project existed long before the competition. It was not invented as a last-minute Build Week concept.
Since October 2024, we have built nearly two years of recorded human-AI interaction across multiple generations of GPT models. This gives the project something that cannot be created during a single hackathon: historical depth.
We have already:
produced dozens of extended AI conversation episodes; preserved a public audio archive; observed the transition between different GPT generations; identified repeated pronunciation and voice failures; documented corrections made by a human during live interaction; sent selected audio-bug reports to OpenAI Support; developed an initial taxonomy of pronunciation and audio problems; created a community around long-form AI testing; demonstrated that real conversations can reveal problems that short benchmarks miss.
We are especially proud of the project’s continuity. A single test shows how one model behaves on one day. Our archive shows how voice AI changes across months and model generations.
We are also proud that the project brings together technology, language, culture, broadcasting, and human perception. It treats the human listener not as background noise, but as an essential part of AI evaluation.
What we learned
We learned that a correct answer is not always a successful voice answer.
Pronunciation, pacing, stress, pauses, tone, and rhythm strongly influence whether a response feels trustworthy and understandable. A technically correct sentence can sound wrong, confusing, or even change meaning when spoken poorly.
We also learned that long conversations reveal different weaknesses than short prompts. A model may perform well during the first several exchanges but later begin repeating itself, losing context, changing style, or failing to apply a correction consistently.
Human corrections are valuable research signals. When a listener interrupts and says that a name, number, or word was pronounced incorrectly, that moment contains more information than a simple positive or negative rating. It shows the original failure, the expected behavior, the correction process, and whether the model successfully recovers.
We learned that model progress is not perfectly linear. A newer model may improve in naturalness while introducing a new instability elsewhere. Longitudinal testing is therefore important because it preserves both improvements and regressions.
We also learned that automatic evaluation alone is not enough. Voice quality is partly technical and partly perceptual. The strongest system combines automated analysis with human linguistic judgment.
Most importantly, we learned that real users continuously produce valuable evaluation data, but those signals often disappear. Our project is designed to preserve those moments and convert them into reusable tests.
What's next for God, Are You There?
The next step is to transform our historical archive into a working voice-model observatory.
We plan to:
process and structure the complete archive; create at least one thousand annotated audio fragments; build at least three hundred confirmed, reproducible test cases; develop a dedicated Russian pronunciation benchmark; compare the same prompts across several GPT generations; create a public dashboard showing improvements and regressions; allow developers to replay individual evaluation cases; add human review and disagreement-resolution tools; support multilingual pronunciation testing; publish a transparent evaluation methodology; create English-language documentation; release a demonstration dataset where privacy and licensing permit; produce regular reports on the evolution of voice-model quality.
We also want to build a self-improving evaluation loop. When a human reviewer corrects the system, that correction should become a new test case. Future models can then be checked automatically against the same problem.
With further development, God, Are You There? can become an independent historical record of voice AI evolution and a practical tool for developers building conversational agents, educational systems, accessibility products, navigation tools, media applications, and multilingual voice interfaces.
Our long-term ambition is simple:
Every new generation of voice AI should be able to demonstrate not only that it is more intelligent, but that it listens more carefully, speaks more clearly, and learns more reliably from human correction.
Project Proposal
God, Are You There? A Longitudinal Laboratory for Voice AI
Important clarification regarding the $100,000
OpenAI Build Week accepts project submissions until July 21, 2026, at 5:00 PM Pacific Time. Participants are expected to submit a working project built with Codex and GPT-5.6, along with a project description, a demonstration video of less than three minutes, a repository containing setup instructions, and the identifier of the primary Codex session.
Projects are evaluated based on technical implementation, design quality, potential impact, and originality.
The $100,000 represents the total competition prize pool, rather than a guaranteed payment to a single winner. Therefore, our proposal should distinguish between two separate objectives:
Objective One: Compete in OpenAI Build Week and qualify for an official competition prize.
Objective Two: Present a justified twelve-month development budget of $100,000 for expanding the project beyond the initial competition prototype.
Project Summary
God, Are You There? is a long-term independent project examining how artificial intelligence speaks, reasons, changes its communication style, and responds to corrections during extended interaction with a human participant.
The project has been running since October 2024. By July 2026, we had accumulated almost two years of regular recorded interaction with several generations of GPT models, beginning with GPT-4.
Unlike conventional benchmarks, in which a model receives a short list of prepared questions, our project studies artificial intelligence in a live, unpredictable, and culturally rich conversational environment.
The recorded episodes include extended discussions about history, philosophy, culture, technology, psychology, everyday life, language, human behavior, and the development of artificial intelligence itself.
Recordings and project materials are available here:
Project audio archive on Castbox
Project community and publications on Telegram
The Castbox archive contains the ongoing audio series Testing God, including Episodes 55, 56, 57, 58, and 59, as well as a wider archive of earlier recordings. The Telegram community provides project updates, discussions, announcements, and links to high-quality versions of the episodes.
What We Have Already Accomplished
Since October 2024, the project team has conducted dozens of extended voice sessions with GPT models.
We have tested:
- the pronunciation of Russian words, names, surnames, and geographical names;
- the reading of numbers, dates, abbreviations, monetary amounts, and complex expressions;
- word stress and accent placement;
- intonation in long and structurally complex sentences;
- voice stability when switching between languages;
- repetitions, omissions, and distortions of individual words;
- the model’s ability to maintain a topic during a long conversation;
- changes in tone after a correction or clarification;
- reactions to contradictory or evolving instructions;
- the ability to separate verified facts from artistic interpretation;
- model behavior during a recorded live-style broadcast;
- the model’s ability to recover after making an error;
- differences between successive generations of GPT models.
During these experiments, we identified pronunciation errors and other audio-related failures. According to the project team, selected audio examples and descriptions of these problems were submitted to OpenAI Support.
Until now, most of this material has existed primarily as an audio series.
The next stage is to transform this archive into a structured voice evaluation system that can be used by researchers, developers, educators, voice-application designers, localization specialists, and quality-assurance teams.
The Problem We Are Solving
Most traditional AI benchmarks measure whether a written answer is factually or logically correct.
However, a voice model can produce a correct written response while still:
- placing stress on the wrong syllable;
- mispronouncing a personal name;
- reading a date or number incorrectly;
- losing the intended intonation of a sentence;
- producing unnatural pauses;
- switching incorrectly between Russian and English;
- changing tempo, pitch, or vocal style in the middle of an answer;
- failing to understand a user’s correction;
- repeating a sentence that has already been spoken;
- changing the meaning of a correct text through incorrect emphasis.
For voice-based artificial intelligence, these are not merely cosmetic defects.
In education, journalism, navigation, customer support, accessibility services, translation, healthcare communication, and applications for elderly users, an incorrect pronunciation or misleading intonation can create confusion, damage trust, or communicate inaccurate information.
As voice interaction becomes a primary way of communicating with AI, systematic and reproducible voice testing becomes increasingly important.
What We Will Build
We will create a working platform under the provisional international title:
Are You There, AI?
A Longitudinal Voice Model Observatory
The Russian-language identity of the project will remain:
God, Are You There?
A Longitudinal Observatory for Voice AI
The platform will convert our accumulated recordings and future conversations into structured and reproducible evaluation cases.
Module One: The Digital Archive
All existing project recordings will be collected in a unified database.
Each relevant fragment will include:
- recording date;
- episode number;
- model version;
- language;
- conversation topic;
- original user prompt;
- written model response;
- spoken model response;
- identified error;
- user correction;
- the model’s subsequent reaction;
- severity level;
- human verification status.
This will transform the project from a collection of media recordings into a structured research corpus.
Module Two: Automated Transcription and Segmentation
Audio recordings will be divided automatically into individual user and model turns.
The system will generate:
- verbatim transcripts;
- timestamps;
- speaker separation;
- segmentation into short testable fragments;
- identification of words that may contain pronunciation errors;
- comparison between intended text and produced audio;
- detection of pauses, repetitions, interruptions, and incomplete phrases.
For live experiments, the system will support real-time transcription. Historical recordings will be processed in batches.
Module Three: A Voice Error Taxonomy
We will develop a dedicated classification system for voice-model failures.
The principal categories will include:
Phonetic errors
Incorrect pronunciation of individual sounds, syllables, or sound combinations.
Stress and accent errors
Incorrect placement of stress within a word.
Numerical reading errors
Incorrect pronunciation of dates, times, monetary amounts, percentages, measurements, and ordinal numbers.
Proper-name errors
Distortion of personal names, surnames, company names, book titles, films, songs, and geographical names.
Multilingual errors
Failures occurring when the model switches between languages or reads foreign-language words within a Russian sentence.
Intonation errors
Incorrect emphasis that changes the logical or emotional meaning of a sentence.
Rhythm and pacing errors
Excessive pauses, sudden acceleration, unnatural slowing, mechanical delivery, or broken sentence rhythm.
Semantic audio errors
Cases in which the written answer is correct, but the spoken version changes or obscures its meaning.
Recovery errors
Cases in which the model accepts a correction but subsequently repeats the same pronunciation defect.
Voice instability
Cases in which identical text is pronounced differently across repeated attempts.
Each confirmed issue will include:
- a short audio clip;
- the expected pronunciation;
- the actual pronunciation;
- a transcript;
- an error category;
- a severity score;
- reproduction instructions;
- the model and date on which it occurred.
Module Four: Comparative Testing Across GPT Generations
One of the project’s greatest strengths is its duration.
We will reuse identical or equivalent prompts across multiple model generations in order to determine:
- which errors disappeared;
- which errors remained;
- which new errors appeared;
- whether pronunciation improved;
- whether long-context stability improved;
- whether models became better at accepting corrections;
- whether speech became more natural;
- whether multilingual switching improved;
- whether long conversations became more coherent;
- whether the model’s voice became more consistent.
The result will not merely be a static leaderboard.
It will be a historical map of the evolution of voice-based artificial intelligence.
How GPT-5.6 and Codex Will Be Used
The GPT-5.6 model family can support different layers of the project according to the complexity and volume of each task.
GPT-5.6 Sol
Sol will be used for the most difficult analytical work:
- comparing behavior across model generations;
- developing and refining the error taxonomy;
- reviewing ambiguous cases;
- analyzing long conversations;
- identifying recurring failure patterns;
- preparing research reports;
- generating new evaluation hypotheses;
- examining whether a correction genuinely changed model behavior.
GPT-5.6 Terra
Terra will support daily production work:
- classifying audio fragments;
- generating concise descriptions;
- applying thematic labels;
- comparing transcripts with intended text;
- preparing draft reports;
- identifying similar errors in the archive;
- organizing evaluation cases.
GPT-5.6 Luna
Luna will be used for high-volume and clearly defined operations:
- renaming files;
- creating metadata;
- preliminary sorting;
- validating file formats;
- identifying duplicates;
- processing large numbers of short audio fragments;
- preparing material for deeper analysis.
Real-Time Voice Models
Real-time voice models will support live comparative sessions:
The same question is presented to several model versions.
The spoken answers are recorded.
The system identifies differences.
A human reviewer confirms or rejects suspected errors.
The result is automatically added to the evaluation catalog.
Codex
Codex will serve as the engineering core of the project.
It will be used to build:
- the audio-upload application;
- the test-case database;
- the experiment replay system;
- the model-comparison dashboard;
- automated validation checks;
- reproducible report exports;
- the public demonstration version;
- technical documentation;
- security and reliability tests;
- a repository that can be reviewed and run by competition judges.
The Core Innovation
The innovation is not simply that we record conversations with artificial intelligence.
We are building a complete evaluation loop:
Live conversation → error detection → audio-fragment extraction → classification → repeat testing → cross-model comparison → human verification → reproducible report
An ordinary listener may notice an unusual pronunciation and forget it several minutes later.
Our platform turns that moment into a permanent test case that can be repeated, evaluated, and used to improve future generations of voice models.
What We Will Present for OpenAI Build Week
For the competition, we will not claim to have completed a massive research platform that exists only as an idea.
We will present a focused, functional prototype.
The competition version will include:
Uploading an audio episode.
Automatically generating a transcript.
Dividing the recording into short segments.
Selecting a suspected voice error.
Assigning an error category.
Comparing the spoken output with the intended text.
Generating a structured test-case card.
Repeating the prompt with a newer model.
Displaying a before-and-after comparison.
Exporting a reproducible evaluation report.
The prototype will demonstrate that our historical archive can be transformed into a practical developer and research tool.
Why the Project Fits the Competition Criteria
Technical Implementation
This is not a simple chatbot or a thin interface built around a single prompt.
The project combines:
- audio processing;
- transcription;
- structured data storage;
- model-based evaluation;
- repeatable experiments;
- human review;
- longitudinal model comparison;
- report generation.
Product Design
The user will see a clear evaluation card showing:
- what the model was expected to say;
- what it actually said;
- where the error occurred;
- how serious the error was;
- whether the problem reappeared in a newer model;
- whether the model corrected itself successfully.
Potential Impact
The platform may be useful to:
- developers of voice agents;
- creators of educational applications;
- researchers studying human-AI interaction;
- Russian-language specialists;
- localization teams;
- media organizations;
- customer-service teams;
- navigation-product developers;
- accessibility-product developers;
- users who rely primarily on voice interfaces.
Originality
The project already has an extensive history of real-world interaction.
We are not inventing a fictional research story solely for a competition. The competition would serve as the moment when an existing two-year audio archive is transformed into an engineering and research product.
Why This Project Matters for the Future
Voice is becoming an independent interface for interacting with artificial intelligence.
A user should not need to look at a screen every time the model reads a surname, medication name, address, date, legal term, or instruction.
A voice system must be understandable, stable, testable, and trustworthy.
Developing such systems requires more than large laboratory datasets. It also requires long-term observations of real interaction:
- how people correct a model;
- which errors are most noticeable to listeners;
- which failures persist after correction;
- how a model behaves during a long conversation;
- how quality changes after an update;
- which improvements are visible to ordinary users;
- where automatic metrics disagree with human perception;
- how trust develops or collapses during repeated interaction.
Our archive makes it possible to study voice AI not only in a controlled laboratory, but in a living conversational and cultural environment.
Why We Need a $100,000 Development Budget
The requested budget should be presented as a twelve-month development plan following the competition prototype, rather than as the value of a single Build Week prize.
Technical Development: $32,000
This funding will support:
- backend development;
- frontend development;
- database architecture;
- model integrations;
- experiment replay infrastructure;
- automated tests;
- report export;
- version management;
- deployment and maintenance.
Audio Archive Processing and Annotation: $12,000
This will cover:
- preparing historical recordings;
- dividing episodes into individual fragments;
- reviewing transcripts;
- adding timestamps;
- identifying duplicate material;
- creating structured metadata;
- manually validating important cases.
API Usage, Computing, and Data Storage: $15,000
This funding will support:
- audio processing;
- comparative model testing;
- storage of original recordings;
- backups;
- operation of the public demonstration platform;
- large-scale archive processing;
- experimentation with new models.
Linguistic Expertise: $10,000
Professional language review will be required for:
- verifying word stress;
- creating reference pronunciations;
- assessing Russian phonetics;
- evaluating names and multilingual fragments;
- developing a severity scale;
- resolving disputed pronunciation cases.
Research and Methodology: $9,000
This will support:
- creation of the benchmark dataset;
- development of evaluation criteria;
- statistical and qualitative analysis;
- preparation of research publications;
- reproducibility testing;
- documentation of experimental procedures.
Design and Visualization: $8,000
This funding will be used for:
- the comparison interface;
- the historical model timeline;
- visualizations of error categories;
- a public improvement dashboard;
- mobile adaptation;
- accessible presentation of audio findings.
Privacy, Security, and Legal Preparation: $6,000
This will cover:
- rules for processing voice recordings;
- participant consent procedures;
- removal of personal information;
- data-retention policies;
- licensing review;
- access protection;
- responsible publication standards.
Demonstration and Educational Materials: $5,000
This will support:
- short educational videos;
- an English-language version of the platform;
- technical documentation;
- public demonstrations;
- materials for developers and researchers;
- project-presentation materials.
Contingency Reserve: $3,000
This reserve will cover:
- unexpected technical costs;
- changes in API pricing;
- additional testing;
- emergency fixes;
- critical infrastructure work.
Total requested development budget: $100,000
The precise allocation would be finalized according to the final team structure, taxation requirements, contractor costs, and actual infrastructure expenses.
Measurable Twelve-Month Outcomes
By the end of the funded development period, the project will deliver:
- a structured archive of historical episodes;
- at least 1,000 annotated voice fragments;
- at least 300 confirmed and reproducible evaluation cases;
- a dedicated Russian-pronunciation benchmark;
- comparisons across multiple generations of GPT models;
- a public results dashboard;
- a tool for replaying and repeating tests;
- reproducible report exports;
- a methodology for human evaluation;
- English-language technical documentation;
- an open demonstration dataset wherever legally and ethically permitted;
- regular reports describing changes in voice-model quality.
Submission-Ready Project Description
Project Title
Are You There, AI? A Longitudinal Voice Model Observatory
Short Description
Since October 2024, the project God, Are You There? has conducted extended, recorded conversations with successive generations of GPT models.
During this period, we have built a substantial audio archive and identified errors involving pronunciation, word stress, numerical reading, multilingual switching, intonation, voice stability, and long-context conversational behavior.
According to the project team, selected audio examples and descriptions of these issues were submitted to OpenAI Support.
For OpenAI Build Week, we are transforming this archive into a working developer and research tool.
The platform transcribes recordings, identifies suspected voice errors, classifies them, converts them into reproducible test cases, and repeats them on newer models.
GPT-5.6 is used for analysis, classification, comparison, and hypothesis generation. Codex is used to build the application, evaluation infrastructure, testing workflow, and comparison interface.
Our goal is to create a long-term system that measures not only how intelligently artificial intelligence responds, but also how clearly, accurately, naturally, and consistently it speaks.
Why Now
Voice models can now participate in live conversations, interpret speech, produce spoken responses, and switch between languages in real time.
The next essential step is to create reproducible systems that show:
- where voice models fail;
- how they respond to human correction;
- whether the same error appears again;
- whether a new generation genuinely performs better than the previous one;
- whether technical improvements are noticeable to real users.
Our project already contains the historical material needed to begin answering those questions.
Long-Term Vision
We intend to transform almost two years of recorded human-AI interaction into an independent observatory for the development of voice-based artificial intelligence.
The platform will preserve the history of model evolution and help developers create voice systems that are more understandable, multilingual, reliable, and worthy of human trust.
God, Are You There? should become more than the title of an audio series. It should become the question that every new generation of artificial intelligence answers in its own voice.
Log in or sign up for Devpost to join the conversation.