CubicLM — Private, Local AI for Android & Windows

CubicLM — Private Local AI

CubicLM

Scroll
PRIVATE AI • LOCAL MODELS • NO CLOUD • OFFLINE FIRST • OPEN SOURCE • PRIVATE AI • LOCAL MODELS • NO CLOUD • OFFLINE FIRST • OPEN SOURCE •
ANDROID • WINDOWS • GGUF • LITERT • STABLE DIFFUSION • VULKAN • ANDROID • WINDOWS • GGUF • LITERT • STABLE DIFFUSION • VULKAN •
What We Build
01
LLM Chat
GGUF LiteRT 21 Providers
02
Image Generation
Stable Diffusion FFI Vulkan
03
Code Runner
Sandboxed Syntax Multi-lang
04
100% Offline
No Cloud No Account MIT
05
Multi-Model Pool
Switch Models Per Chat Hive Service
06
Privacy & Backup
App Lock JSON Backup AES-256
07
Battle Arena
4 Models Live Monitor Verdict
Download CubicLM
Download CubicLM
Windows
Windows
Full desktop app: chat, 20+ cloud providers, backup & restore, App Lock, and Model Hub with keyboard shortcuts.
ZIP · v1.12.0 · Windows 10+ · 17 MB
Download for Windows
Android
Android
On-device inference with LiteRT-LM and llama.cpp. Run GGUF models directly on your phone — no cloud needed.
APK · v1.12.0 · Android 9+ · arm64 · ~62 MB
Download for Android
Android9+ · 4 GB+ RAM recommended · free GBs for models (GGUF 0.5–8 GB each)
Windows10+ · WebView2 · cloud models (on-device GGUF needs the Android app)
Why CubicLM
Private by Default
No telemetry, no accounts, no cloud. Your conversations stay on your device with AES-256 encryption. Your data never leaves.
23+ Providers
OpenRouter, OpenAI, Anthropic, Gemini, Groq, DeepSeek, and more. Switch providers per chat with a single tap.
Image Generation
Stable Diffusion 1.5 runs locally via FFI with Vulkan acceleration. Generate images offline on your device.
Code Runner
Sandboxed code execution with syntax highlighting. Run, test, and iterate — all within the chat interface.
Local API Server
Built-in OpenAI-compatible server at port 8080. Connect any tool or app to your local models on your network.
Web Aware
Optional web fetch augments responses with live, cited sources. No API keys needed — just toggle it on.
20+
Cloud Providers
15
Languages
100%
Offline Capable
MIT
Open Source
Deep Dive
Features
Battle Arena
Race up to 4 AI models on a single prompt — same provider or mixed providers. Same-time mode fires all contenders at once; One-by-one mode runs them serially in pick order and even allows your loaded on-device model. A live monitor ranks who finished first, who is fastest (tokens/sec), and who wrote most, then declares an overall winner with a transparent 40/30/30 weighted score. Find it in the Toolkit tab.
4 Contenders · Live Monitor · Verdict
Slide Maker
Describe a topic and the AI designs every slide — title, bullets, visuals, speaker notes. Text-only models reserve proper image boxes; Stable Diffusion renders them. Regenerate any single slide, edit by hand, reorder, then export Markdown, PDF, or standalone HTML. Find it in the Toolkit tab.
AI Deck · Regen · MD/PDF/HTML
CubicWeb Builder
Tell AI what to build — files stream live into the explorer with a live-reloading preview, v0-style. Plan mode, undo history, diff view, responsive viewports, screenshot-to-code, one-click deploy to Vercel/Netlify, plus React/Vite/Next.js dev-server pipeline with validation and crash recovery. Find it in the Toolkit tab.
Live Build · Deploy · Dev Servers
Developer Terminal
A real shell in the app: run commands in the project workspace with streaming output, command history and autocomplete. One-tap installs for AI coding CLIs — Claude Code, OpenCode, Cline and Kilo — then launch them inside your project and interact directly.
Shell · CLI Manager · History
System Logs
When valid code still won't run, CubicWeb System Logs tells you why: structured CW-* diagnostics separate code bugs from missing runtimes and device limits, so the AI stops rewriting good code and shows you the real fix instead.
CW-* Codes · No-loop Fixes
Local AI Inference
Run LLMs on-device via llama.cpp (GGUF) with GPU acceleration through Vulkan / OpenCL, or Google's LiteRT-LM runtime for .litertlm models. Streaming tokens with real-time speed display.
GGUF + LiteRT + GPU
20+ Cloud Providers
OpenRouter, OpenAI, Anthropic, Gemini, DeepSeek, Groq, Mistral, Together AI, and 12+ more. Import fresh models from /models, test every model to see what's really online, auto-hide the dead ones, and auto-sync on your schedule.
Unified Plugin Architecture
Image Generation
On-device Stable Diffusion 1.5 via FFI with Vulkan acceleration. Generate images completely offline — DreamShaper, CyberRealistic, Realistic Vision, and more safetensors models.
SD 1.5 · Vulkan · FFI
Skills
Offline prompt-injection extensions in markdown. No network, no SDK — just text appended to the system prompt. Intelligent per-prompt activation scores and selects the top 2 relevant skills.
5 Built-in · Import · Browse
MCP Server
Connect a single remote MCP server via Streamable HTTP / SSE. Tools are added to OpenAI-compatible cloud requests with automatic tool-call round-tripping and size-capped results.
HTTP / SSE · Bearer Auth
Local API Server
Expose local models as an OpenAI-compatible API on port 8080. Use your downloaded GGUF or LiteRT models from any compatible client on your network — no cloud required.
Port 8080 · OpenAI Format
Web Fetch
Toggle on in chat to automatically fetch and inject web content from URLs in your message. Pages are stripped to clean text, with visible source chips (favicon + domain) shown below responses.
3 Links · 9K Chars · 15s Timeout
Multi-Session Chat
Full chat history with Hive persistence, full-text searchable sidebar, message actions (copy, share, regenerate, branch, edit), revision history, and code blocks with syntax highlighting.
Search · Export · Branch
App Lock
Biometric gate on launch and background-resume — Android biometrics or Windows Hello, with device-PIN fallback. Your conversations stay yours, even on a shared device.
Hello · Fingerprint · PIN
Backup & Restore
Export every conversation to a portable JSON backup and merge-restore it on any device. Pin important chats so they always float to the top of history.
JSON · Pin · Merge
Read Aloud
Listen to any reply with on-device text-to-speech. Locale-matched voices across all 15 languages — tap a bubble to start, tap again to stop.
TTS · 15 Locales
Desktop Power
Keyboard-first on desktop: Ctrl+N new chat, Ctrl+F history search, Ctrl+1–4 tabs — plus a close guard that asks before quitting a running generation or download.
Shortcuts · Close Guard
On-Device & Image Gen
Supported Models
On-device GGUF / LiteRT / SD models run on Android. On Windows, chat via 20+ cloud providers — or save models to your Downloads folder for later.
LiteRT-LM — Google's on-device runtime
Model Size Description
Qwen3-0.6B LiteRT 586 MB Smallest chat model for low-RAM phones
Qwen2.5-1.5B Instruct LiteRT 1.49 GB Balanced int8 quantized chat model
DeepSeek-R1-Distill-Qwen-1.5B LiteRT 1.71 GB Reasoning-focused distilled model
Gemma 4 E2B Instruct LiteRT 2.46 GB Google Gemma vision + chat
Gemma 4 E4B Instruct LiteRT 3.40 GB Highest quality LiteRT option
GGUF — llama.cpp with GPU acceleration
Model Size Description
Kimi Moonlight 16B-A3B GGUF 7.1 GB MoE, 3B active parameters — Q3_K_S
Qwen2.5-3B Instruct GGUF 2.1 GB Best mobile speed/quality — Q4_K_M
Qwen2-VL-2B GGUF 1.5 GB Vision-capable — Q4_K_M
Phi-3.5 Mini GGUF 2.2 GB Microsoft reasoning model — Q4_K_M
Gemma 2 2B GGUF 1.71 GB Google lightweight chat — Q4_K_M
Llama-3.2-3B Uncensored GGUF 2.1 GB Unrestricted assistant
Llama-3.2-1B Instruct GGUF 0.8 GB Ultra-lightweight entry model
Image Generation — Stable Diffusion 1.5
Model Size Description
DreamShaper 8 LCM SD 1.5 2.0 GB Fast 4-step generation
CyberRealistic V8 FP16 SD 1.5 2.0 GB Photorealistic, uncensored
Realistic Vision V5.1 FP16 SD 1.5 2.0 GB Popular portrait/scene model
AbsoluteReality 1.8.1 SD 1.5 2.0 GB General-purpose photorealistic
AnyLoRA SD 1.5 2.0 GB Anime / stylized
Multi-Provider Cloud
Cloud Providers
A unified plugin architecture — each provider implements the CloudProvider interface with auto-fetched model lists, free tier tagging, and per-provider company filters.
OpenRouterFree
Hugging Face
xKiroFree
TokenRouter
AgentRouter
OrcaRouter
APInex
OpenAI
Anthropic
Google Gemini
DeepSeek
Z.AI
Groq
Mistral AI
Together AI
xAI Grok
Perplexity
Cerebras
Fireworks AI
Cohere
NVIDIA NIM
Stability AI
Kimi
Custom
Custom
Kimi
Stability AI
NVIDIA NIM
Cohere
Fireworks AI
Cerebras
Perplexity
xAI Grok
Together AI
Mistral AI
Groq
Z.AI
DeepSeek
Google Gemini
Anthropic
OpenAI
TokenRouter
APInex
OrcaRouter
AgentRouter
xKiroFree
Hugging Face
OpenRouterFree
Under the Hood
Tech Stack
Framework
Flutter 3.x Flutter 3.x
Languages
Dart Dart Kotlin Kotlin C++ C++
State Management
GetX GetX
Local Storage
Hive Hive SecureStorage SecureStorage
Inference Engines
llama.cpp llama.cpp LiteRT-LM LiteRT-LM SD FFI SD FFI
Cloud & Services
Firebase Core Firebase Core Crashlytics Crashlytics Messaging Messaging
Codebase
Project Structure
lib/
├──dartmain.dart# App entry point
├──core/
│ ├──dartcolors.dart# App color palette
│ ├──dartconstants.dart# Settings keys, model catalog, API endpoints
│ ├──dartroutes.dart# Route definitions
│ ├──darttheme.dart# Light/dark theme
│ ├──dartdesign_tokens.dart# Warm palette, spacing, typography
│ ├──dartlanguages.dart# 15 supported languages
│ └──dartapp_translations.dart# ~160 keys × 15 languages
├──models/
│ ├──dartai_model.dart# AI model data class
│ ├──dartchat_message.dart# Chat message with revisions, sources, skills
│ ├──dartchat_session.dart# Chat session model
│ ├──dartweb_source.dart# Web source (url/domain/favicon/title)
│ ├──darttask_model.dart# Automated task model
│ ├──dartnotification_entry.dart# Model-switch history
│ └──dartskill_model.dart# Skill (name/description/content/enabled)
├──controllers/
│ ├──dartchat_controller.dart# Chat logic, streaming, web/skill tracking
│ ├──dartcloud_model_controller.dart
│ ├──darthome_controller.dart# Tab navigation, model resume
│ ├──dartmodel_controller.dart# Model download/import
│ ├──dartserver_controller.dart# Local API server
│ ├──dartsettings_controller.dart# App settings, locale, system prompt
│ └──darttask_controller.dart# Automated task execution
├──services/
│ ├──dartcloud_service.dart# Multi-provider cloud API
│ ├──cloud/# 20+ provider plugins
│ │ ├──dartcloud_provider.dart
│ │ ├──dartcloud_provider_registry.dart
│ │ └──providers/# openai, anthropic, google, deepseek, groq, etc.
│ ├──dartinference_service.dart# Cross-platform inference orchestrator
│ ├──dartinference_android.dart# Android llama.cpp / LiteRT bridge
│ ├──dartopenai_server_service.dart
│ ├──dartdownload_native.dart# Resumable streaming downloader
│ ├──dartdownload_service.dart# Download orchestrator
│ ├──darthive_service.dart# Hive persistence, full-text search
│ ├──dartnotification_history_service.dart
│ ├──skills/# Skill registry, injector, GitHub/URL sources
│ ├──mcp/# MCP config, connection, registry
│ ├──dartdevice_info_service.dart# RAM/tier + SoC/GPU detection
│ ├──dartweb_fetch_service.dart# URL fetching → clean text + sources
│ ├──dartexecution_service.dart# Task execution engine
│ ├──dartdocument_extractor_service.dart
│ ├──dartlocal_image_service.dart# Stable Diffusion inference
│ ├──dartsd_isolate_processor.dart
│ ├──dartimage_generation_notification_service.dart
│ ├──dartapp_log_service.dart# Logging with crash pattern detection
│ └──dartcrash_reporting_service.dart
├──views/
│ ├──dartsplash_view.dart# 1380ms shimmer + fade-out
│ ├──dartonboarding_view.dart# 3-page PageView
│ ├──darthome_view.dart# Main navigation scaffold
│ ├──dartchat_view.dart# Chat interface
│ ├──dartmodel_view.dart# Model Hub — 4-way toggle
│ ├──dartexplore_skills_mcp_tabs.dart
│ ├──dartserver_view.dart# Nodes page (Node + Config)
│ ├──dartsettings_view.dart# Config sections
│ ├──dartapp_settings_view.dart# Theme, typography, Thinking Orbs, Language
│ ├──dartlanguage_picker_view.dart
│ ├──dartabout_view.dart
│ ├──dartnotification_history_view.dart
│ ├──dartlog_view.dart# System diagnostics
│ └──darttask_view.dart
├──widgets/
│ ├──dartchat_bubble.dart# Message bubble with actions + revisions
│ ├──dartcode_block.dart# Syntax-highlighted code
│ ├──dartmodel_switcher_sheet.dart
│ ├──dartattachment_preview.dart
│ ├──dartimage_viewer.dart
│ ├──dartthought_disclosure.dart
│ ├──dartthinking_orb.dart# 3D particle sphere animation
│ └──darttyping_indicator.dart
├──ffi/
│ └──dartsd_ffi_bindings.dart# FFI bindings for SD native lib
├──utils/
│ ├──dartapp_snackbar.dart# Top spring-animated toast
│ └──dartthought_parser.dart# <thought> tag parser
└──shared/# re-export of root shared/
├──constants/
│ └──dartplatform_links.dart
└──theme/
└──darttokens.dart
shared/# root single source (for future shells)
├──constants/
│ └──dartplatform_links.dart
└──theme/
└──darttokens.dart
windows/# Flutter Windows shell
├──runner/
├──CMakeLists.txt
└──updater_config.json
web/# Flutter Web shell
├──index.html
└──manifest.json
scripts/
├──build-all.ps1
└──build-all.sh
local_plugins/
├──llama_flutter_android/# llama.cpp Flutter plugin
├──flutter_litert_lm/# Google LiteRT-LM Flutter plugin
└──sd_flutter_android/# Stable Diffusion Flutter plugin
android/
├──app/src/main/kotlin/com/cubiclm/app/
│ ├──MainActivity.kt
│ └──ModelDownloadService.kt# Foreground service
└──res/values+drawable/
└──launch_background.xml
assets/
└──skills/# 5 bundled starters
docs/
├──ARCHITECTURE.md
├──PLATFORM_DIFFERENCES.md
├──BUILD_AND_RUN.md
├──PLATFORM_LINKS.md
└──../CHANGELOG.md
Releases
Changelog

v1.12.0+19 2026-09-09

Added

  • Intelligent code editing — ghost text, inline AI edits, diff review, @mentions.
  • Studio preview — browser header, resizable viewport, version timeline.

Fixed

  • Silent model-load death — service declared, native errors surface, kill-proof breadcrumb.
  • Engine isolation — GGUF/LiteRT separated.
Full changelog on GitHub →
Questions
FAQ
Yes — MIT open source, no accounts, no telemetry. Local models are free forever; cloud providers need your own API keys (many have free tiers, tagged Free in the app).
Not yet — on-device llama.cpp / LiteRT / Stable Diffusion are Android-only. Windows gets full chat, 20+ cloud providers, backup, shortcuts and App Lock. A native Windows engine is on the roadmap.
The app itself is ~61 MB. Models range from 0.5 GB (Qwen3-0.6B, runs on 4 GB phones) to 8 GB (7B+ quants). 6 GB+ RAM is recommended for 7B models; keep free GBs for downloads.
Chats live on your device in AES-256 encrypted storage with an optional biometric lock. Cloud mode sends prompts only to the provider you explicitly chose — nothing else leaves the device.
Android updates in one tap from inside the app (About → Check for updates). On Windows, download the new ZIP from this page. The website always points at the latest release.
15 — including Bangla, Hindi, Urdu, Arabic, Chinese, Spanish, French and more. Switch instantly in App Settings; voices and UI follow your choice.
Feedback
Report an Issue
Found a bug or have a suggestion? Help improve CubicLM by reporting it directly on GitHub.

You'll be redirected to GitHub to complete the issue creation (requires GitHub account).

Your AI,
Your Rules

Download Now