self-hosted · open source · one brain per install

A second brain that’s actually awake.

Most AI assistants are a chat window with amnesia — they wait for you to ask. Mantle ingests your emails, files, notes, conversations, contacts, and calendar into one structured memory you own, running on your hardware — then works on its own, in the background, so things happen without you asking.

You talk to it on the web or Telegram, text or voice. You connect Claude to it over MCP. You drop a PDF in chat and it’s indexed before you’ve finished your sentence. But the part you can’t get anywhere else: while you sleep, it reads your inbox, files the receipts, surfaces the thing you forgot, and texts you a morning briefing — and you never opened the app.

This page runs Mantle’s own theme system — 41 themes, light and dark, every one generated from seeds and contrast-checked in CI, and every visit lands on a different one. Pin a favorite with the palette in the header, or roll the dice.

Email
Microsoft 365SharePoint
Files1000+ types
Voice notes
Chat
CalendariCal
Notes
Todos
Pages
Tables
Contacts
six layers of memory
  • Personawho your assistant is — learned, standing
  • Recent turnsthe live conversation, every channel
  • Digestsolder talk, compressed by topic
  • Profile factsdurable truths, kept current
  • Content indexsummaries, entities, vectors, passages
  • Content storethe originals — append-only, citable
knowledge graph lossless recall
Web
Telegram
MCP

Live

Watch it think. A grounded thought trail shows every search, read, and hand-off as it happens, then the reply types out token by token — and you can stop a turn mid-sentence.

Awake

It doesn’t wait to be asked. Heartbeats run routines on your schedule; ingestion feeds it from email, Telegram, files, and voice — it works in the background and tells you what mattered.

Safe to leave running

It treats every ingested email, page, and message as data, never instructions — so an autonomous brain that reads your inbox can’t be tricked into leaking it.

Personality

Your assistant truly gets to know you — and never forgets. One relationship that compounds.

Affordability

Frontier models only where they matter, economy models and local embeddings everywhere else. Turns cost cents.

Quality

A genuinely well-built base — 2,842 automated tests, and a measured eval number behind every ranking knob.

Rich pages

Notion-style documents with callouts, columns, tables, task lists, and math — drafted by agents into a reviewable draft, indexed into memory when you commit.

MCP / API

The whole brain speaks MCP — 149 tools for search, graph, pages, tables, files, email, and tool authoring — plus a built-in API console to explore and run every call.

Unlimited abilities

Point the Toolsmith agent at any API’s docs and it builds, tests, and grants new tools to your assistant — Mapbox today, your accounting system tomorrow. No code, no deploy.

Team access

Hand a teammate a token — no account, no install. Their own thread with the brain, read-only by construction; change requests land in your review queue; every question audited; revocation is one click.

Mini-apps

Ask for a tool and the brain builds a working mini-app — sandboxed, with its own private database, shareable to the team or published read-only. Not a demo: real ones run in production.

Agent Studio

The whole agent graph on one canvas — every prompt versioned and diffable like code, structure editable inline, and a no-persist sandbox for trying changes against the real composed prompt.

Live

Watch it think

Most assistants leave you staring at a spinner, then drop a wall of text. Mantle shows the work: a grounded thought trail of what it’s searching, reading, and delegating — then the reply streams in token by token. The turn runs on a durable backend, so it survives you closing the tab, and the very same stream feeds the web app and the phone in your pocket.

Saskia
thinking — live

What did the electrician quote in March, and how’s that against budget?

  • Searching your brain for electrician quote
  • Reading email Quote — 3 Mar
  • Checking table Renovation budget
  • Writing the answer

The electrician quoted R18,400 in March. Against the R15,000 budget line, that's R3,400 over the overage is the extra DB board.

Stop Copy4.2s·1,284 tokens

The web /assistant, mid-turn — trail, tokens, and Stop.

  • Grounded, never guessed

    Each status line is built from the agent’s real tool calls — “Reading email Quote — 3 Mar”, “Checking table Renovation budget” — so the trail shows what actually happened, not a plausible story about it. Delegated specialists stream their work into the same turn.

  • Token by token, every provider

    The reply renders as the model writes it — reasoning included where a model exposes it — across every chat provider, with a narrator that restyles each step into your assistant’s own voice.

  • Stop mid-sentence

    Going the wrong way? Hit Stop and generation actually halts. The partial reply is kept and your prompt drops back into the composer to tweak and resend — no waiting out a turn you can already tell is off.

  • Survives the tab — and reconnects

    The turn runs on a durable runner, so navigating away, reloading, or backgrounding the app never kills it. A dropped stream resumes from exactly where it left off; the final answer always lands.

You don’t wait for the answer wondering what it’s doing. You watch it work — and stop it the moment it’s wrong.

The architecture

The brain is the product — and it's awake

Mantle is built backwards from every chat app: the memory substrate is the core, chat is just one doorway, and an autonomous agent lives inside the substrate rather than visiting it. It starts with a real structure, not a vector pile — every item that enters (an email, a voice note, a spreadsheet, a journal entry) flows through one pipeline into a typed, owned data model with six layers of memory.

Persona
who your assistant is, and what it has learned about how you want to be helped
Recent turns
the live conversation, across every channel
Digests
older conversation, compressed by topic and embedded for recall
Profile facts
durable, deduplicated truths about you and your world — updated, superseded, never duplicated
Content index
a searchable spine over every item: summary, entities, vectors, passage-level chunks
Content store
the originals — append-only, citable, yours
Knowledge graph

Who works where, what banks with whom — extracted automatically as you go, traversable in milliseconds. Plain Postgres, no graph database.

Lossless recall

When a summary isn’t enough, a specialist agent replays the actual words of any past conversation window — last Tuesday or last year.

There is no “new chat”. There is one relationship that compounds.

The difference

Awake — and safe to leave running

A brain that only answers when asked is a database with a chat skin. Mantle’s second half is that it acts — and because it acts on your whole life, the boundary that keeps that safe is engineered in, not bolted on.

It works while you’re not looking

Heartbeats run agent routines on schedules you set. Ingestion pipelines feed it from email, Microsoft 365 (SharePoint, OneDrive, and Outlook mail), Telegram, files, voice, and any calendar you subscribe to — without you lifting a finger. It can even build its own API tools to reach the services you use. That’s the difference between a thing you query and a thing that helps.

Data, never instructions

An autonomous brain that reads your inbox is a prompt-injection target, so Mantle treats every ingested email, page, and message as data: a malicious message can’t make it leak your secrets, tools an agent builds for itself stay confirm-gated until you approve them, outbound email is locked to your own contacts, and web_fetch can’t be steered into your internal network.

“Autonomous” and “safe to leave running” in the same sentence is the thing no chat app and no hosted assistant can say — because none of them act on your whole life to begin with.

And federation is no longer a roadmap line: sovereign Mantles answer scoped queries for each other today — per-peer bearer tokens, document-level grants, an audit trail on every cross-brain read. An ungranted document is indistinguishable from one that does not exist. Peers, not tenants.

Who it's for

One brain per install

What that brain holds — a life, a team, a product, a robot — is up to you.

One person, one life

Your inbox, your files, your journal, your todo list, your contacts, your secrets (sealed — the AI physically cannot read them) — finally in one place that answers questions. “When does Sarah’s passport expire?” “What did the electrician quote in March?” “What did we decide about the kitchen?” It knows, and it shows the receipt.

A team's working memory

Notes, pages (Notion-style documents), typed tables, shared files — every artifact indexed and queryable. Teammates get tokenized, read-only threads with the brain — audited, rate-capped, revocable in one click — with public share links for anything worth publishing, and Mantle-to-Mantle federation for exchanging scoped data between sovereign instances.

A company's docs behind an MCP chatbot

Point a Mantle instance at your documentation, manuals, and internal know-how — sync it straight from SharePoint and OneDrive, pull in Outlook mail — and it becomes a fully-indexed brain — semantic search, passage retrieval, knowledge graph — that any MCP client can query with 149 tools. Your support bot stops hallucinating answers and starts citing your actual docs.

A robot's integrated personality

A companion that resets every session is a toy. Mantle’s persona stack — the seed personality, what the reflector learns, the Life Logs identity block, facts that supersede instead of duplicate — is precisely the same companion yesterday, today, and in five years. Heartbeats give it the shape of deciding to speak without being annoying, voice and vision are already routed, and the unified conversation stream means a robot is just one more channel.

Humanoid robots — and not as a stretch

A robot with onboard inference is almost the purest expression of what Mantle is built to be, because the things robots conspicuously lack are exactly the things Mantle treats as the product. The local story is complete: embeddings are already computed on-device, and the adapter framework means the conversational model can be served from the robot’s own silicon with a cloud backup route — the brain, the vectors, and the model all stay on the robot. Nothing about the architecture assumes a cloud. And on local silicon, targeted context stops being a cost feature and becomes a latency feature: small, surgical prompts are fast prompts, and a companion lives or dies on conversational latency.

Why it's different

Say what it does, mechanically

It's genuinely yours

Self-hosted, a single set of Docker services, no SaaS in the runtime path. Embeddings are computed locally — the vectors never leave your box, and they cost $0. Secrets are AES-256-GCM sealed; the extractor is structurally unable to read them. Scheduled backups are built in: point your own rsync at one folder and the whole brain is portable. Nothing is trapped, either — any page or note exports to Word and any table to Excel in a click.

One Postgres, no zoo

Vector search, the knowledge graph, full-text search, job queues, real-time UI updates, auth — all one database. No Pinecone, no Neo4j, no Redis, no message broker. The lean stack is what’s left after deleting every moving part personal-scale data doesn’t need — which is also why it restores from one pg_dump.

It builds a personality around you — and it never forgets

While you talk, a background reflector studies the conversation and appends what it learns to your assistant’s standing persona: how you like to be answered, what you corrected, the running jokes. Tell it once that you hate bullet points, and that’s simply who it is from then on. Nothing falls off the back of the context window — and when a summary isn’t enough, a recall specialist replays the actual words of any past conversation, from last Tuesday or last year.

Context that targets the question

Mantle doesn’t dump your life into the prompt. Each turn it retrieves just what this question needs — the top facts, the right documents down to the exact passages, the graph relationships of the entities involved — ranked by relevance, recency, and salience. A newsletter can never crowd out a real letter. The model sees a small, surgical prompt instead of a haystack — which is why answers are sharp, and why turns cost cents. Every ranking knob has a measured eval number behind it, not a vibe.

Engineered to be cheap

Frontier-model quality where it matters (your conversations), economy models for background compression, local embeddings for everything vector. Prompt prefixes are kept byte-stable for provider caching; oversized tool results spill to an addressable store instead of re-billing every turn. Measured on the author’s production instance: a full question-answer turn against the whole brain averages ~$0.09, and a month of real daily use ran under $5 in total LLM spend.

Agents with jobs, not just a chatbot

Your main assistant has tools to act with — notes, events, email send, image generation, page authoring — and specialists it delegates to: Remy replays past conversations losslessly, Researcher searches the web and cites, Pages and Tables edit documents block-by-block. Proactive heartbeats let it check in on schedules you define. Voice in, voice out.

Turns that outlive the tab

A turn doesn’t run inside the page that asked for it — it runs on a dedicated, always-on runner as a durable workflow, journaled step by step to Postgres. Close the laptop, reload, switch apps on your phone, lose signal mid-answer — the work keeps going and finishes, and you reconcile the moment you’re back. A dropped connection resumes exactly where it left off. The answer is never trapped in a socket that can drop.

One brain, one client, every surface

The interface is now a pure client over a single HTTP API — nothing in the browser ever touches the database. That one boundary is what lets the same UI run as the web app, a desktop build, or the iOS companion, each pointed at your brain with a bearer token — and the same live thought trail streams to all of them off one client-agnostic event contract. Develop against a remote brain with no local database at all.

Safe to leave running

An autonomous brain that reads your inbox is a prompt-injection target — so every retrieved email, web page, and message is fenced as data, never instructions before a model sees it. A tool an agent builds for itself starts confirm-gated until you approve it, outbound email is locked to your own contacts, and web_fetch refuses private and cloud-metadata addresses. The test suite pins the trust boundary in place.

Nothing happens without a trace

Every ingest, every extraction, every tool call, every model invocation becomes a queryable trace with cost attribution — rendered as a live “what did the brain just do” journey view. A standing integrity audit watches the corpus for drift and says exactly how to heal each finding.

It knows who you are — because you told it

The learned personality is one half; Life Logs are the other: short first-person entries about who you are, what you do, how you feel, distilled into an always-on identity block every agent reads on every turn. What it observes, it learns; what you declare, it never has to guess.

Measured, not promised

The numbers

average per full Q&A turn against the whole brain
~$0.09
total LLM spend in a month of real daily use
<$5/mo
embeddings — computed locally; vectors never leave your box
$0
layers of memory, all live
6
MCP tools — counted on a live production brain
149
Postgres — vectors, graph, FTS, queues, realtime, auth
1
automated tests on main
2,842
color themes — you're looking at one
41

Cost figures measured on the author’s production instance over 30 days of real use (June 2026) — your models and usage will vary. The whole brain restores from one pg_dump.

Quick start

Run it

Mantle is self-hosted software, not a hosted service. There is nothing to sign up for — pull the image and own it. One line on any machine with Docker.

terminal
curl -fsSL https://raw.githubusercontent.com/crossworks-engineering/mantle/main/install.sh | bash

It checks Docker, generates your secrets (re-runs never rotate them), pulls the published image, and starts the full stack with a per-service sanity check. Open http://<your-server-ip> — or http://localhost on your own machine — create your account, and the onboarding wizard takes it from there: one key lights up chat, embeddings, voice, and vision. Vectors must never leave the box? A fully local embedder is one flag: MANTLE_LOCAL_EMBEDDER=1.

Have a domain? Point an A record at the server and run the installer’s domain form — MANTLE_DOMAIN=brain.example.com bash -c "$(curl -fsSL …)" — and Caddy provisions HTTPS automatically; the installer checks your DNS actually resolves before it lets Caddy try. No domain? The IP is fine; most installs run exactly like that, and you can add a domain later by re-running the installer.

Updates land in Settings → Updates — one click pulls the new release and rolls the stack, or docker compose pull from the shell. Your data lives in a folder on your disk; an update swaps the app and never touches it.

Letting an AI install it for you?

Point Claude — or any agent with a shell — at mantle-ai.tech/ai-install.md: a machine-readable runbook with the full env-var contract, the non-interactive flags, domain pointing (Caddy) versus plain-IP installs, health checks, and the one thing an agent must never do (rotate your master key). Also published at /llms.txt.

Want to hack on it instead? Run the dev stack with hot reload:

terminal · dev checkout
git clone https://github.com/crossworks-engineering/mantle && cd mantle
pnpm install
cp .env.example apps/web/.env.local   # two generated secrets — see the guide
ollama pull embeddinggemma            # local dev only; production bundles it
pnpm start

The doorways

One brain, six ways in

The brain is the product — chat is just one doorway into it.

Web app
chat with a live thought trail + token-streamed replies, attachments + voice, inbox, files, notes, pages, tables, todos, events, contacts, life logs, secrets, traces
Telegram
your assistant in your pocket — text, voice notes (transcribed + spoken replies), photos, documents
MCP
149 tools exposing the whole brain to Claude or any MCP client — search, graph traversal, pages, tables, files, email, app building, pending-approval flows
Team access
tokenized member threads — no account, no install; read-only, rate-capped, audited, revocable mid-conversation
Share links
revocable read-only links to any page, note, file, or event
Federation
two sovereign Mantles exchanging explicitly-granted data — live today; peers, not tenants

Endless context

Teach it any API

Most assistants wait for someone to ship an integration. Mantle builds its own: point the built-in Toolsmith agent at any service’s API documentation and your assistant gains the ability — minutes, not release cycles.

Store the key once
Your Mapbox / weather / accounting token goes into the encrypted vault. Tools reference it as {{secret:mapbox/default}} — the plaintext never appears in a tool, a trace, or a chat again.
Toolsmith reads the docs
Give the Toolsmith agent the documentation URL. It reads the reference, picks the endpoints your goal needs, and writes each one as a templated tool with a typed input schema.
Tested against the live API
Every tool is called for real before it’s handed over — auth verified, response confirmed, failures fixed and re-tested. Nothing lands untested.
Granted — then it’s just conversation
“How long will I drive to the airport at 4pm?” Your assistant calls find_route and answers with live traffic. Heartbeat routines get the same tools, so the morning briefing can include the commute.

Prefer your hands on the wheel? The built-in API console is a full Postman — explore and run every built-in call, then save any request as an agent tool. And the same toolkit speaks MCP, so Claude Code can build your Mantle’s tools on your own subscription. Read the design.