-
-
Automated regression test suite passing 15/15 unit tests with 100% green checkmarks for ONNX parsing and determinism.
-
Executive compiler dashboard displaying INT8 quantization, Flash/SRAM allocation, 0 B heap, and latency estimates.
-
Multi-target hardware matrix comparing memory fit and execution speed across STM32H7, ESP32-S3, and RP2040 microcontrollers.
-
Physical SRAM memory arena using interval graph coloring for mathematically verified zero-collision buffer reuse.
-
Topological Shannon IR computation graph showing layer connections, tensor shapes, and activation buffer lifetimes.
-
Symmetric INT8 quantization studio delivering 4x memory compression, 49.9 dB SQNR, and per-layer scale calibration.
-
Standalone C99 header emission (shannon_model.h) with 4-way loop unrolling, static BSS arena, and zero heap allocations.
-
Bare-metal C firmware script (main.c) running the zero-malloc inference loop with DMA sensor input and cycle counters.
-
Backend compiler engine (memory_planner.py) running greedy interval graph coloring for zero-allocation MISRA-C safety.
-
Shannon static compiler pipeline: from ONNX model input to zero-malloc standalone C firmware for edge microcontrollers.
Inspiration
The idea for Shannon came from a problem that kept bothering us while working with TinyML:
A neural network can be completely fine in Python and still be practically unusable on a microcontroller.
On a laptop, we don't usually think twice about allocating another tensor, using a runtime, or letting a framework manage memory for us. On a microcontroller with only a few tens of kilobytes of SRAM, those assumptions stop working.
We wanted to understand what actually happens between:
"I trained a model."
and
"This model is running on the chip."
What we found was that a surprisingly large part of the problem is not the neural network itself. It is the deployment machinery around it — runtime overhead, tensor allocation, memory lifetime management, quantization, operator implementation, and hardware-specific constraints.
So instead of building another inference runtime, we decided to take the opposite approach:
What if we removed the runtime and compiled the model directly for the hardware?
That became Shannon.
What We Built
Shannon is a static TinyML compiler for severely memory-constrained microcontrollers.
The basic workflow is:
ONNX Model → Intermediate Representation → INT8 Quantization → Tensor Lifetime Analysis → SRAM Arena Allocation → Hardware Fit Check → C/C++ Code Generation
The important part is that Shannon does not simply convert a model into C.
It tries to understand the model as a memory and execution graph.
For every tensor, we track its lifetime through the graph. If two tensors are never alive at the same time, they do not need separate physical SRAM.
We therefore treat SRAM allocation as an interval-allocation problem.
Overlapping tensor lifetimes cannot share memory, while non-overlapping lifetimes can reuse the same offset.
This allows Shannon to construct a single contiguous activation arena instead of allocating every intermediate tensor independently.
The generated firmware uses static memory:
static uint8_t shannon_tensor_arena[ARENA_SIZE]
rather than malloc() / calloc() / realloc() / free().
That makes the generated memory behaviour deterministic before the firmware is even executed.
The Compiler Pipeline
1. Model Parsing
We implemented actual ONNX parsing rather than relying on a simplified demonstration graph.
The parser reads the ONNX protobuf representation, extracts operators, tensor dimensions and initializers, validates the graph schema and converts it into Shannon's internal representation.
Unsupported or malformed operators are rejected explicitly.
We deliberately avoided silent fallbacks because a compiler that quietly changes what the model means is much worse than one that simply tells you it cannot compile the model.
2. Symmetric INT8 Quantization
Weights and activation scaling are converted to signed INT8 representation.
For a tensor with maximum absolute value (M), Shannon uses:
[ S = \frac{M}{127}, \qquad Z = 0 ]
where (S) is the scale and (Z) is the zero point.
We also calculate quantization quality metrics rather than only reporting the new tensor size.
These include:
- Mean Squared Error (MSE)
- Signal-to-Quantization-Noise Ratio (SQNR)
- Vector cosine similarity
This lets us see the numerical cost of reducing precision instead of treating quantization as a magic "compress" button.
3. SRAM Lifetime Planning
This is probably the part of Shannon we are most proud of.
Suppose tensor A is required by layers 1–3 and tensor B is required by layers 4–6.
There is no reason for A and B to occupy different SRAM regions permanently.
Shannon represents tensor lifetimes as intervals and uses greedy interval-based allocation to reuse physical offsets whenever lifetimes do not overlap.
The result is a contiguous SRAM arena containing:
Input → Reused Activation Buffers → Output
with explicit offsets for each tensor.
We then verify the allocation for collisions and calculate the exact peak activation requirement.
4. Static Hardware-Fit Verification
Once the model is compiled, Shannon checks whether the generated representation actually fits the selected target.
We track at least two hard resources:
- Flash required by quantized weights
- SRAM required by the activation arena
If the model exceeds the target's available memory, the compiler should expose that constraint instead of pretending the deployment is valid.
5. C/C++ Generation
Finally, Shannon emits a standalone C header containing the model representation, weights, activation arena and inference implementation.
The generated kernels use INT8 arithmetic and loop-unrolled execution where applicable.
There is no inference interpreter sitting underneath the generated model.
The goal is that the generated file is something an embedded developer can actually take into a firmware project.
Why We Chose a Compiler Instead of Another Runtime
Initially, it is tempting to solve TinyML deployment by putting another abstraction layer between the model and the chip.
But abstraction has a cost.
A general-purpose runtime needs to understand graphs, operators, tensors and memory dynamically because it has to support many different models.
Shannon knows the model before deployment.
That changes the problem.
Once the graph is known at compile time, many decisions that would normally happen at runtime can be made beforehand:
- tensor locations
- activation lifetimes
- memory reuse
- quantized weight representation
- operator execution order
- buffer sizes
- hardware memory limits
We therefore trade some runtime flexibility for deterministic and much smaller deployment code.
For the class of fixed TinyML workloads we are targeting, we think that is a worthwhile trade.
Building It
The compiler core is written in Python, while the developer-facing workstation is built with React, TypeScript and Vite.
The repository is intentionally split into several layers.
The Python compiler contains the IR, ONNX parser, quantizer, memory planner and C code generator.
We also implemented a TypeScript version of the core compiler algorithms so the browser workstation can perform the same type of analysis locally.
The backend exposes the compiler through FastAPI, while the frontend presents the compilation process as a hardware-oriented workstation rather than a generic AI dashboard.
We also added firmware starter projects for different microcontroller families, including ESP32, RP2040, STM32, Teensy 4.1 and nRF52840.
The Part That Took the Most Time
The difficult part was not getting a neural network to produce an answer.
The difficult part was making sure the compiler was telling the truth.
It is very easy to make a TinyML demo look impressive by displaying estimated latency, memory numbers and "optimized" percentages.
We did not want to do that.
For example, Shannon's latency values are explicitly treated as static cycle estimates, not measurements from an oscilloscope.
Similarly, the compiler distinguishes between memory that is actually planned and telemetry that is only estimated.
We also built a regression suite covering things such as graph differences, memory collisions, dynamic allocation checks, ONNX parsing and deterministic code generation.
At the moment, the repository contains 15 compiler and regression tests.
What We Learned
The biggest thing we learned is that TinyML is as much a systems problem as it is a machine-learning problem.
A model can be accurate enough and still fail deployment because its activation tensors do not fit in SRAM.
A model can fit in Flash and still be impractical because of runtime overhead.
A quantized model can be smaller but introduce unacceptable numerical error.
And generated code can compile successfully while still having a fundamentally bad memory layout.
So our thinking changed from:
"How do we run this neural network on a microcontroller?"
to:
"What can we prove about this neural network before it ever reaches the microcontroller?"
That shift is essentially what Shannon is about.
Challenges
The hardest engineering challenge was connecting three very different worlds:
Machine Learning → Compiler Design → Embedded Hardware
Each one has different assumptions.
The ML graph describes computation.
The compiler needs to turn that computation into deterministic operations and memory locations.
The hardware imposes hard limits on SRAM, Flash, alignment and execution resources.
Getting these layers to agree was significantly harder than implementing any individual feature.
Memory planning was particularly interesting because a seemingly small change in the graph can change tensor lifetimes and therefore the entire SRAM layout.
We also had to make sure that optimizations did not silently change model behaviour.
That is why validation and regression testing became part of the compiler itself rather than something we added at the end.
What Shannon Means to Us
We don't see Shannon as "AI that magically optimizes models."
It is much more specific than that.
It is an attempt to make neural-network deployment behave more like compilation.
Give the compiler a known graph.
Give it a hardware target.
Give it hard constraints.
Then make as many deployment decisions as possible before runtime.
The end goal is simple:
A developer should be able to start with a trained model and end with hardware-oriented C/C++ without manually rebuilding the memory and deployment layer for every model.
There is still a lot we want to improve — more operators, better kernel selection, deeper hardware-specific optimization, more aggressive quantization strategies and eventually measured silicon feedback.
But Shannon already gives us something we wanted from the beginning:
A concrete path from model → memory plan → generated code → microcontroller.

Log in or sign up for Devpost to join the conversation.