The Technology Powering Soulbits

An open-source AI engine ecosystem built with local-first setups in mind, backed by scalable cloud with maximum focus on privacy.

SOULBITS SOFTWARE

A complete ecosystem - built for local execution

Soulbits App

The Harmony AI App brings the full companion experience to your pocket. Available as custom iOS IPA and Android APK builds, it works completely offline — your companion goes everywhere you go, with no cloud dependency required.

Full offline capability — no internet connection requiredCustom iOS (IPA) and Android (APK) buildsComplete companion experience: chat, voice, and personalitySyncs with Harmony Link sessions across devices

Documentation

iOS (IPA)
Android (APK)

k

Screenshot Placeholder — Soulbits AI App UI

Local Quickstart Suite

Get the entire Harmony AI stack running locally with a single docker-compose setup. Includes Harmony Link, our Text-Generation-WebUI fork, and the Harmony Speech Engine.

Docker Hub

Soulvoice Engine

Harmony Speech is a high-performance AI Speech Engine that enables faster-than-realtime speech generation, zero-shot voice cloning, and speech recognition using state-of-the-art open-source AI voice technology. It provides OpenAI-style APIs for seamless integration.

OpenAI-style APIs for Text-to-Speech, Speech-to-Text, Voice Conversion, and Speech EmbeddingMulti-model and parallel model processing supportToolchain-based request processing with internal re-routing between modelsExtensible architecture for adding future open-source AI speech models

Supported Model Architectures

OpenVoice V1
OpenVoice V2 / MeloTTS
OpenAI Whisper (FasterWhisper)
Harmony Speech V1
Screenshot Placeholder — Speech Engine UI

Local Quickstart Suite

Get the entire Harmony AI stack running locally with a single docker-compose setup. Includes Harmony Link, our Text-Generation-WebUI fork, and the Harmony Speech Engine.

Docker Hub
ARCHITECTURE FLOW

How it works

1
User Input (Text or Voice) → Capture and Processing via Soulbits App
2
Soulbits Engine retrieves historical context, RAG backstories, and system instructions
3
Payload is routed to the configured AI backend (Local PC hardware via Soulbits AI Link or secure Cloud)
4
LLM generates response → Processed via Soulvoice Engine for emotive voice synthesis
5
Real-time streaming output (Audio, Text, and Live2D/3D animation triggers) delivered to the user interface

Supported Backends & Models

Total architectural freedom. Choose the model stack that fits your privacy goals and hardware capabilities.

Large Language Models (LLM)

Llama 3, Mistral, Gemma, Phi-3, and OpenAI-compatible API endpoints (Run locally via Ollama, LM Studio, or KoboldAI).

Text-To-Speech (TTS)

StyleTTS2, XTTSv2, AllTalk, ElevenLabs, and cloud-hosted neural providers.

Speech-To-Text (STT)

OpenAI Whisper (Local faster-whisper implementations) and browser-native web speech APIs.

Vision & Image Gen

Stable Diffusion (XL/v1.5) via Automatic1111/ComfyUI web APIs for dynamic companion expressions and image generation.

Download Our Applications

Get the official mobile clients to take your companion with you wherever you go.

Looking for local desktop orchestration? Check out the Cloud or Docs tab for raw build scripts and server binaries.