Skip to content
SS
Press the power button on the remote to start the Shikhar Sisodia workspace.
FinderFileEdit
shikhar/projects/local-edge-ai-sdk.project

In progress2026

Local Edge-AI SDK

An offline-first SDK for running small language models on-device, from Python, C and Flutter.

v0.3.1 · verified on desktop and in CI · not yet run on a physical phone

year
2026
role
Builder
status
In progress

01what it is

What it does.

A Python pipeline that compresses Hugging Face models to GGUF with llama.cpp, a native C runtime around llama.cpp with token streaming and cancellation, and Dart/Flutter bindings for Android and iOS. No API key, no cloud inference, no telemetry.

Verified so far: Phi-3 Mini Q4_K_M generated and run on Windows, the real runtime on Intel macOS, Android packaging and an emulator run, and the iOS simulator. Not yet: inference on a physical phone, and there is no Neural Engine or NPU support.

02the product

The live product.

local-edge-ai-sdk-umber.vercel.app
The Local Edge-AI SDK project site: AI that stays on your device, three panels for local by design, compression pipeline and mobile SDK, and a note that inference runs locally in the SDK, not on the site

The project site. A status page, not a demo: inference runs in the SDK on a device, and the page says so.

03engineering

How it is built.

Model conversion, a native C ABI, Python ctypes and Dart FFI bindings with contract tests, a Flutter plugin with Android and iOS packaging, and model-free CI on Linux, macOS and Windows.

  • Python
  • C/C++
  • Dart
  • Flutter
  • llama.cpp

04where it stands

Still being built.

In development; the project site is live and the SDK is at version 0.3.1.