Infrastructure for AI that stays on-device

Build local AI without rebuilding the stack.

CactusBrain gives developers SDKs and optimized models for private, offline AI on mobile and edge devices. Powered by cellm, our open-source local inference engine.

Inference runs locally. App data is not sent to CactusBrain.

cactusbrain.stack
03

YOUR APPProduct experience

02

CACTUSBRAIN SDKsVision · Text · more

01

Powered by cellmLocal inference runtime

InferenceOn-device
NetworkNot required
Data sent0 bytes

The products developers use. One engine underneath.

CactusBrain brings SDKs, optimized models, and distribution together. cellm powers local execution without becoming another integration developers have to manage.

01 / Developer platform

CactusBrain

Documentation, optimized models, accounts, projects, entitlements, and automatic encrypted model delivery.

discover → integrate → ship
02 / Developer products

Vision · Text

Focused CactusBrain SDKs with task APIs, validation, and models developers can ship.

SDK → optimized model → result
03 / Engine underneath

cellm

The open-source runtime that loads compatible models and executes inference locally.

CPU · Metal · local execution

Boundary: CactusBrain does not proxy inference, and the SDKs do not duplicate cellm runtime logic.

Purpose-built interfaces over the same local runtime.

Each SDK owns its workflow, validation, and developer experience. cellm remains the execution layer underneath.

01Developer Preview

CactusBrain Vision

On-device visual identification for apps, with matching that stays local.

  • Image embeddings
  • Custom visual collections
  • Multiple reference images
  • Similarity search
  • Unknown-object rejection
  • Fully local matching
Explore CactusBrain Vision
02Available

CactusBrain Text

On-device language model inference for private text features.

  • Local text generation
  • Qwen 2.5 and LFM packages
  • CPU and Metal execution
  • Offline after model download
View compatible models

The engine underneath CactusBrain.

cellm is our open-source inference runtime for running optimized AI models directly on mobile and edge devices. CactusBrain SDKs add product workflows and stable APIs on top; inference stays in one inspectable engine.

RustOpen sourceQuantized inferenceCPUMetalC ABI / FFICLI
View cellm on GitHub URL pending

Packages with explicit compatibility and licensing.

Vision packages will appear only after a compatible cellm runtime adapter and model artifact are available.

ModelTaskTargetsLicenseAccess
Alibaba CloudQwen 2.5stable
texttext_generation
iOS / macOSCPU + Metal
Apache 2.0Permitted
Entitlement requiredSign in →
Liquid AILFMstable
texttext_generation
iOS / macOSCPU + Metal
LFM Open License v1.0Revenue threshold applies
Entitlement requiredSign in →

From catalog to a shipped app.

  1. 01Create an account

    Access the developer workspace and available SDKs.

  2. 02Create a project

    Choose a target platform and connect model access to an app.

  3. 03Ship it

    Assign an entitled model package. The SDK delivers and loads it on the device when your feature first runs.

Open a developer workspace