Inspiration

Bloom's famous Two Sigma Problem showed the extraordinary potential of one-to-one tutoring: when a student has a good personal tutor who understands them, is patient with them, adapts to how they learn, identifies where they are struggling, and guides them individually, their learning outcomes can improve dramatically.

The problem is that this does not scale.

It is simply impossible to give every student a truly great personal teacher—especially across communities where good educational resources, strong teachers, reliable internet access, and even computing power may already be limited.

As an aspiring educator myself, I have always thought about the absence of practical, real-world applicable learning in African education. And as a machine-learning enthusiast, I kept thinking about the same question:

What if technology could give every student that kind of deeply personal educational companion?

That idea became Muta.

Muta is our attempt to bring the promise behind Bloom's Two Sigma Problem closer to reality: an Education AI companion that can meet a student at their level, patiently teach them, adapt to them, and remain available even when the internet is not.

And importantly, we are building it for the realities of African students, not only for classrooms with perfect connectivity and expensive hardware.


What it does

Muta is an offline-first educational AI application built around an AI model with strong capabilities in high school STEM education, while also extending into arts and commerce.

A student can ask Muta a question naturally and receive an explanation, continue asking follow-up questions, work through mathematical and scientific reasoning, view interactive visualizations, use audio, and communicate in more than 30 supported languages.

But something is extremely important to us:

Muta is not the model.

The AI model is one component of Muta, and we deliberately designed the system so that the underlying model can be changed as better models become available.

Our long-term moat is instead the deeply integrated educational intelligence layer around that model.

We envision Muta understanding not only a student's latest question but also increasingly understanding the relationships between:

student → teacher → parent → school → curriculum → assessment → learning outcomes → best career. Muta will be the perfect education-to-career platform guiding every African student.

Over time, Muta should understand what a student knows, what they misunderstand, what they have forgotten, how they learn best, what their teacher is currently teaching, what their curriculum expects from them, and what they should learn next.

The underlying AI model may be replaced.

The accumulated education graph, learner context, institutional relationships, workflows, and learning intelligence cannot be replaced nearly as easily.

That is the Muta we are building toward.


How we built it

Muta is composed of five major systems:

  1. The Muta Frontend
  2. The Muta Backend
  3. The Inference Engine
  4. The AI Model
  5. The Muta Fleet Manager

We designed Muta as real installable desktop software rather than simply another web interface connected permanently to a cloud model.

The frontend is built with React, HTML, and CSS, while the backend is written in Python with FastAPI, providing an asynchronous interface between the application and Muta’s local services. Application data is persisted locally using SQLite, with PostgreSQL support for larger deployments.

The most demanding engineering work was context and user-state management. Muta must reconstruct multi-turn conversations while fitting system prompts, tutoring modes, personas, learner preferences, learning-twin data, retrieved resources, images, web context, reasoning, and responses within a finite context window. This required:

  • intelligent history trimming without modifying stored conversations;
  • preservation and resumption of interrupted streams;
  • persistent conversations, messages, attachments, citations, preferences, and mastery data;
  • strict per-user data isolation;
  • authentication, sessions, permissions, CSRF protection, and secure uploads;
  • reliable and complete account deletion.

The backend also:

  • coordinates concurrent learners through fair request queues, cancellation, reconnection, and replayable streams;
  • supervises and restarts the local llama.cpp inference engine;
  • dynamically manages RAM, KV cache, context capacity, and model slots;
  • safely degrades vision, speech, or session capacity under memory pressure;
  • integrates RAG, PDF processing, image understanding, speech, mathematical verification, and sandboxed tools behind a stable API.

For distribution, the Python backend is packaged for each operating system using PyInstaller, while Tauri bundles the complete desktop application for our three primary targets:

  • Windows
  • macOS
  • Linux

The result is a system designed to run securely, privately, and reliably without installation complexity, a dedicated GPU, or continuous internet access.

For local AI inference, Muta supports llama.cpp, allowing the model to execute directly on commodity CPUs. Rather than carrying unnecessary parts of the larger native stack, we compile the components Muta needs and use FFmpeg for our media-processing requirements.

And because running AI locally means that every CPU, battery, and amount of RAM is different, we built Muta to understand the machine it is running on.

Muta can inspect the available system resources and adjust how much work it safely performs. Its memory-management system determines how many simultaneous conversations the machine can support and how much context can safely be allocated to each one.

We also built Muta Power Optimization. When a student is running on battery, Muta can enter Eco Mode, bounding automatic reasoning and using shorter responses where appropriate to reduce unnecessary computation while preserving full reasoning budgets when the task genuinely requires them.

Muta also includes Host Mode, allowing one computer running Muta to privately serve other users on the same local network. This is particularly important to where we believe Muta can go: a school should not necessarily need a powerful computer—or an internet connection—for every single student before local AI becomes useful.

Finally, we built the Muta Fleet.

Muta itself is capable of operating offline, but whenever an installation reaches the internet, the Fleet allows us to receive pseudonymous application heartbeats, understand deployment health, track versions and platform characteristics, improve the product from real-world usage, and provide the foundation for safely distributing future updates.

The cloud therefore supports Muta.

It does not define whether Muta works.


Challenges we ran into

Our first challenge was surprisingly fundamental:

What exactly does "low-resource" mean in education?

At first, it is tempting to define the problem only as a lack of textbooks, teachers, internet access, or powerful computers.

But we quickly discovered that the problem is much wider.

A student can have a textbook and still lack someone who can patiently explain it. A school can have internet access but not have enough bandwidth for every student to continuously use a cloud AI service. A student can have a laptop, but it may have only a few gigabytes of usable RAM and no dedicated GPU. A learner can understand English and still understand a difficult concept much better when it is explained in the language they think in.

So our challenge became not simply building an AI model.

It became building an educational system that can operate across different levels of hardware, connectivity, language, learning ability, and educational support.

The second major challenge was the engineering trade-off between intelligence and resources.

Every improvement in reasoning, context size, parallel conversations, visual interaction, or model capability has a computational cost. We constantly had to measure memory usage, CPU utilization, inference speed, model size, and response quality and ask:

How much intelligence can we deliver inside the machine the student already owns?

That question influenced everything from our inference runtime and model choices to Eco Mode, memory limits, and how many conversations Muta allows to generate simultaneously.

The third challenge was turning an AI experiment into an actual product.

Supporting one development computer is very different from shipping software across Windows, Linux, and macOS. Native dependencies, application packaging, local inference, media handling, updates, operating-system differences, and resource detection all had to work together.

We wanted Muta to be something someone could actually install and use—not merely something that worked on our own machines.


Accomplishments that we're proud of

In the short span of roughly two months—and an almost unreasonable number of sleepless nights—we went from an idea to software that real people could install and use.

We are especially proud that Muta is not simply a browser demo whose intelligence disappears when the internet connection disappears.

Muta runs its AI locally on the user's own machine.

We successfully built and shipped Muta across the three major desktop operating systems: Windows, Linux, and macOS. That means the same educational experience can reach students using very different computers without requiring them to buy specialized AI hardware.

We built interactive visualizations because sometimes an explanation should not only be read. A student learning vectors, geometry, physics, or another spatial concept should be able to see and interact with what is being explained.

We added audio because education should not be restricted to typing and reading alone.

And we added support for more than 30 languages, with particular attention given to African languages, because we believe intelligence should not suddenly become less accessible because a student's strongest language happens to be Igbo, Hausa, Yoruba, Kiswahili, isiZulu, or another African language.

We are also proud of the systems behind the visible product: local inference, dynamic resource management, parallel conversations, power-aware reasoning, offline operation, local-network hosting, cross-platform packaging, and the Muta Fleet that allows us to maintain a growing installation base without making those installations dependent on the cloud.

Most importantly, we have already put Muta in the hands of early users.

Their feedback has validated the problem we are trying to solve while simultaneously showing us just how much work remains.

Our mission remains simple:

Meet every African student at their level.


What we learned

The most fascinating thing we learned was how much a product changes once real people begin using it.

In the beginning, there was naturally a fear that external feedback would undermine what we had built. Instead, the opposite happened.

Our early testers found bugs our late nights could not find. They asked questions we had never considered. They used features differently from how we expected. Some of their frustrations forced us to reconsider assumptions that had seemed completely reasonable while developing Muta ourselves.

And with every round of feedback, our confidence increased—not because users told us everything was perfect, but because they showed us that the problem was real enough to keep solving.

That process has already taken Muta through three versions.

We also learned that constraints can create features.

Limited battery life led us to think seriously about power-aware reasoning.

Limited RAM forced us to build adaptive memory management.

Limited internet access reinforced our decision to make local inference fundamental rather than optional.

The possibility that several students may need access while only one suitable machine is available led us toward local-network Host Mode.

The realities we originally saw as limitations increasingly became part of the architecture of Muta itself.


What's next for Muta

Muta is not stopping with the student.

Our goal is much larger than building a chatbot that answers homework questions.

Remember our moat:

Muta is the deeply integrated educational intelligence layer connecting students, teachers, parents, institutions, and curriculum—accumulating context, workflows, relationships, and learning intelligence over time.

The next stage is to begin turning that vision into infrastructure.

We want Muta to build a persistent understanding of each learner: the concepts they have mastered, the misconceptions they repeatedly encounter, the explanations that work for them, what they are likely to forget, and what they should learn next.

We want teachers to understand where an entire class is struggling without having to individually inspect hundreds of conversations.

We want the teacher's lesson, the student's learning, the assessment, the curriculum, and eventually the wider institution to stop existing as disconnected pieces of information.

We want them to become one connected educational graph.

And we’re building it to work without internet access—over a local area network. It’s a difficult engineering challenge we’re choosing deliberately, so Muta can reach even the lowest-resourced learners and institutions.

Moreover, because the AI model underneath Muta is deliberately replaceable, improvements in AI should strengthen this system rather than make the entire product obsolete.

A new model can make Muta smarter.

It should not replace everything Muta has learned about the learner and their educational journey.

This is also where Muta connects to the broader vision behind UDO.

Muta begins with perhaps the most fundamental part of that journey: helping a young person learn, understand, develop their abilities, and discover what they are capable of.

Our ambition is not merely to give African students access to AI.

It is to build technology that understands the educational journey around them well enough to genuinely help them move forward.

From understanding today's lesson, to mastering a subject, to discovering their strengths, to eventually navigating what comes after school.

Muta starts with the lesson. The vision goes much further.

Built With

Share this project:

Updates

posted an update —

Update: Since accuracy remains our highest priority for an educational model, we will only carry forward the vocabulary-pruning optimization. It improved deployment efficiency and reduced model size without any measured loss in accuracy. Our updated deployment model is therefore Muta-Tutor-Qwen2.5-1.5B-Q4_K_M-vocab32k.gguf, which preserves the capability of our selected Muta Tutor while being smaller and faster. It can be found here.

Log in or sign up for Devpost to join the conversation.

posted an update —

Muta ADTC 2026 — Progress Update

From Gate 1 submission to our current model

At Gate 1, we submitted a fine-tuned Qwen3.5-0.8B Q4_0 model for Muta. Our optimization at the time strongly favored the ADTC combined objective: accuracy, performance, and efficiency on an 8 GB RAM, CPU-only machine. The submitted model achieved a strong efficiency profile and successfully passed the initial screening.

Gate 1 metric Score
Accuracy & Quality 39.20
Performance 56.00
Efficiency 90.65
Total 54.53

The judges' feedback, however, exposed the main trade-off we had made: we had optimized too aggressively for speed and efficiency at the expense of reliability. The model could produce convincing-looking answers while making early arithmetic mistakes, mixing units or currencies, giving weak scientific analogies, or propagating an incorrect first step through the rest of the solution.

For an educational product, this is not acceptable. From that point onward, we redefined our priority: accuracy means critical-thinking reliability—the ability to reason correctly, follow instructions, remain internally consistent, correct misconceptions, and teach safely.


Gate 2: choosing a stronger intelligence core

We revisited the stronger model we had already identified during Gate 1: our fine-tuned Muta Tutor Qwen2.5-1.5B Q4_K_M. Although it was larger and slower under the scalar profiler, it was consistently stronger on STEM reasoning.

In our matched evaluation, Qwen2.5-1.5B improved ARC-Easy accuracy from 70.2% to 77.8% compared with the submitted Qwen3.5-0.8B, and performed better on several of the judges' mathematics, science, and explanation tasks.

We then widened the search instead of assuming Qwen2.5 was automatically the final answer. We tested a broader field of small models—including MiniCPM5, LFM2.5, Qwen3/3.5 variants, VibeThinker, Falcon-H1, and OpenReasoning Nemotron—across:

  • 500 ARC-Easy questions
  • 100 custom STEM prompts
  • all Gate 1 judge prompts
  • CPU speed and memory measurements

Across that field, Muta Tutor Qwen2.5-1.5B remained the strongest overall development control, with:

Metric Muta Tutor Qwen2.5-1.5B
ARC-Easy 77.8%
Full STEM passes 57/100
Core-correct STEM 80/100
Gate 1 re-test 47/100
Completed judge answers 10/10
AVX2 score proxy 84.1383

This became our quality-first model going forward.


Improving accuracy: what we tried

We then asked whether the selected Muta Tutor could be made substantially more accurate through further fine-tuning.

We built a 2.5M-row STEM data warehouse spanning mathematics, physics, chemistry, biology, integrated science, tutoring styles, and WAEC/WASSCE material. Rather than train blindly on all 2.5M rows, we created a balanced 300,350-row training sample to control subject imbalance, repetition, leakage, and low-quality rows.

We ran:

  1. Eight 20K-row BF16 LoRA hyperparameter pilots
  2. Three full-data training lineages
  3. A separate science + multi-turn tutoring fine-tune
  4. Held-out evaluations on science, practical STEM, tutoring quality, judge prompts, and STEM MC

The key result was consistent: lower validation loss did not automatically produce a better tutor. Several new checkpoints looked better during training but regressed on fresh evaluation, tutoring behavior, or generalization.

Our final matched comparison still favored the incumbent Muta Tutor:

Test Incumbent Muta Best science challenger (P2)
Held-out science 83.565% 83.565%
Practical 2,000 41.325 39.850
Tutor quality 22/64 12/64
Judges 57/100 50/100
STEM MC 32/50 32/50

This was an important outcome: our earlier model-selection and fine-tuning work had produced a stronger checkpoint than we initially realized. Further training increasingly showed diminishing returns and regression/catastrophic-forgetting risk.

Decision: retain Muta Tutor Qwen2.5-1.5B as the quality-first model.


Optimization: recovering speed without giving up the model

Once the quality-first model was fixed, we changed the question again:

How much can we compress and accelerate the selected Muta Tutor without destroying the capability that made us choose it?

We explored vocabulary pruning, layer pruning, knowledge distillation, MoE conversion, FFN width pruning, Q4_0 quantization, and quantization-aware training (QAT).

The strongest compression path was:

  • vocabulary reduced to 32K
  • depth reduced to 26 layers
  • FFN width reduced to 7168
  • distillation on verified teacher data
  • QAT under simulated Q4_0 noise

This produced our current high-efficiency deployment variant:

refine-qat100-Q4_0.gguf
1.05B parameters · ~593 MB · 15.51 tok/s · ~706 MB peak RAM

Compared with the published Muta Tutor, scalar decode speed increased from roughly 5.5 → 15.5 tok/s, while peak RAM fell from roughly 1.1 GB → 0.7 GB.

However, the compressed model still sacrifices some of the tutoring and reasoning quality of the original Muta Tutor. We therefore treat it as a high-efficiency deployment variant, not yet as the new quality winner.


Where we are now

We now have two clearly defined model tracks:

Track Model Purpose
Quality-first Muta Tutor Qwen2.5-1.5B Q4_K_M Best current tutoring/reasoning model
Efficiency-first refine-qat100-Q4_0 ~593 MB, ~15.5 tok/s scalar CPU deployment variant

The main lesson from Gate 1 to now is that Muta cannot be optimized by chasing a single metric. Accuracy, tutoring quality, throughput, memory, quantization format, CPU kernels, data quality, and training objective all interact.

Our current direction is therefore deliberate: keep the original Muta Tutor as the intelligence benchmark, and only promote a compressed or further fine-tuned model when it matches that benchmark on fresh reasoning and tutoring evaluations—not merely on training loss or speed.

Log in or sign up for Devpost to join the conversation.

Submission history