Inspiration

AI applications on edge devices often face a trade-off between local execution and cloud execution. Local inference can reduce latency and improve privacy, but it may consume significant device resources. Cloud inference can provide more capability, but depends on network availability and adds communication overhead.

We wanted to build a system that can intelligently decide where an AI workload should run based on the current device and network conditions.

What it does

ARMEdge is an AI inference orchestration platform for ARM-based edge devices such as Snapdragon-powered devices.

It continuously considers factors such as battery level, CPU usage, RAM usage, temperature when available, network conditions, and workload characteristics.

Based on these conditions, ARMEdge dynamically routes an AI request:

Local: Phi-3 through Ollama when the device is suitable for local inference. Cloud: Gemini when local execution is unsuitable or unavailable.

ARMEdge also records the routing decision, selected model, reason, confidence, device condition, network state, and inference latency.

How we built it

We built ARMEdge using a Python FastAPI backend and a React/Vite frontend. The backend handles device telemetry, workload analysis, routing decisions, and AI inference.

For local inference, we use Phi-3 through Ollama. For cloud inference, we use the Gemini API.

The frontend provides a dashboard where users can submit AI workloads and observe the selected execution path, telemetry, reasoning, and latency.

We containerized the application using Docker and deployed the cloud-accessible components using Render

Challenges we ran into

One of the main challenges was designing a routing strategy that considers multiple changing conditions instead of simply choosing one model permanently.

We also had to handle the difference between local Docker networking and cloud deployment. During deployment, we configured the frontend and backend as separate services and updated the API routing accordingly.

Another challenge was ensuring that the system could clearly explain why a particular execution path was selected.

Accomplishments that we're proud of

We are proud that ARMEdge goes beyond simple AI model selection and focuses on inference placement.

The same type of workload can receive different execution decisions when device or network conditions change. The system also provides measurable telemetry and inference latency to support those decisions.

We successfully containerized and deployed the backend and frontend as separate services.

What we learned

We learned how edge AI systems must balance performance, resource availability, network conditions, latency, and cloud capabilities.

We also gained practical experience with FastAPI, React, Docker, Ollama, Gemini API integration, GitHub, and cloud deployment.

Most importantly, we learned that effective edge AI is not only about choosing a powerful model—it is also about deciding where and when the model should run.

What's next for ARMEdge

Our future improvements include:

More advanced workload-aware routing. Better prediction of battery and thermal impact. Support for additional local AI models. More ARM/Snapdragon-specific hardware telemetry. Adaptive routing based on historical performance. Federated learning and privacy-aware execution. More edge devices and heterogeneous hardware support.

Built With

Share this project:

Updates

Submission history