SLOPSHOPPER

homie-persona-cognition

Host-bound cognitive lifecycle events for The Homie.

newguardpromptprocess
A shopper browsing a rack in a slop shop
README

TaskChad OS

A self-hosted cognitive agent OS and personal AI assistant with persistent memory, realtime voice, multi-agent orchestration, browser control, Telegram/Discord, and operator-controlled factories.

License: MIT Latest release Python 3.12+ Channels: 6

It hears you live, not voice notes. Talk your agents through real work out loud, the way ChatGPT Voice does on the Codex app, except it's your own second brain and it's open source. Steer a running agent mid-flight by voice, and it remembers the whole conversation after you hang up. Two doors: the dashboard, or /talk join in your Discord voice channel.

It can even install its own voice. Your second brain runs the talk-mode-setup skill, checks your keys, wires the sidecar, and defaults to your Codex subscription so there is no per-minute meter.

Under the voice is a real cognitive OS, not a chat wrapper: a 9-layer cognition stack, multi-agent orchestration over a dependency-tracked convoy graph, and a memory that forms beliefs and flags contradictions instead of agreeing by default. It is built to push back, not just please. It runs on Claude, Codex, Gemini, or an OpenAI-compatible backend and reaches you through Telegram, Slack, WhatsApp, the web, and the CLI. The regression suite is concentrated on the stateful boundaries where agent systems usually fail; see Proof: Tests + Operator Loops.

The Dark Factory

TaskChad OS can assemble specialized personas with isolated memory, scoped tools, scheduled work, team rooms, approval queues, and durable receipts. That makes it a dark-factory toolkit for agent work: the repetitive machinery can keep moving while the operator retains the switches, approvals, and evidence.

"Dark factory" does not mean unbounded autonomy. External writes stay default-denied, coding dispatch stays operator-controlled, and a generated artifact is never treated as pushed, deployed, published, or proven without its corresponding receipt.

Factory map

Factory surfaceWhat it assemblesStart here
Persona factoryIdentity, memory, tools, readiness, and learningPersona Blueprints
Repository factoryBounded issue-to-worktree coding dispatchArchon Repo Dispatch
Visual factoryGrounded image concepts and validated prompt packsImage Node Factory
Authority factoryEvidence-gated SEO/GEO pages and release wavesTokenMax Authority Stack
Client-site factoryBranded, testable service-business websitesClient Site Factory

Lineage + Provenance

TaskChad OS is the original public export of this framework, maintained by the TaskChad OS contributors. It evolved from Cole and the TheHomie Community's Claude Code Second Brain workshop, then grew into an identity-first agent OS with its own memory, orchestration, multi-channel ingress, Operating Room, and desktop surfaces.

OpenClaw, Hermes Agent, OpenSouls, and ClaudeClaw are credited as ecosystem influences. TaskChad OS is an independent project and is not affiliated with, sponsored by, or endorsed by those projects. See NOTICE.md and AUTHORS.md.


Archived Preview

The Homie v0.1.0-alpha.1 dashboard product tour

This 45-second alpha-era product tour shows dashboard, Desktop Stack controls, Mobile Access, Browser Viewer, Work Queue, Convoy, Operating Room, and clean shutdown proof. It is preserved as historical evidence, not a current full-product walkthrough. Realtime voice, quick-agent steering, Runs, and the later factory surfaces are documented in the latest release and the operator manual.

Full-quality MP4 is attached on the v0.1.0-alpha.1 release.


Proof: Tests + Operator Loops

The suite lives in .claude/scripts/tests/. Instead of publishing a count that drifts whenever a generated export changes, TaskChad OS points to the actual coverage surfaces and the commands that reproduce them:

SubsystemRepresentative coverage
Orchestrationtest_orchestration_api.py, test_executor_boundary.py, team and mailbox suites
Cognition + memoryliving-self acts, test_living_memory.py, test_episodes.py, test_session_brief.py, recall and belief suites
Talk Mode voicesession, runs, steering, flush, Discord receive, and debrief suites
Runtime + lane routingselection, lane, provider adapter, diagnostics, and quiet-JSON suites
Memory pipelinesreflection, weekly synthesis, dream, indexing, and recall suites
ObservabilityLangfuse trace-shape and failure-visibility suites

On top of unit coverage, the framework is exercised through operator-loop and smoke testing, not just happy-path assertions:

  • Fresh public Windows install smoke from a clean clone — install, setup check, real CLI chat, Desktop launch, route checks, clean shutdown.
  • Real CLI chat proof after setup, using the same runtime path as channels.
  • Desktop package and portable-app smokes with Python/Hono lifecycle startup, route checks, and clean shutdown.
  • Dashboard route smokes across /mission, /chat, /mobile, /browser, /work, /convoy, and /teams.
  • Langfuse trace validation for the message lifecycle: session lookup, process detection, recall, region assembly, runtime execution where supported, and post-response.
  • Sanitizer/export leak checks before public release so private vault data, local tokens, and machine-specific proof artifacts stay out of the framework.

Run the targeted suite for the surface you changed and use uv run pytest tests/ -q for the broad local regression. Proof boundaries (what is not yet claimed) are listed under Current Proof Boundaries.


Quick Install

# Linux/macOS/WSL
curl -sSL https://raw.githubusercontent.com/TheSmokeDev/taskchad-os/master/install.sh | bash
# Windows PowerShell
irm https://raw.githubusercontent.com/TheSmokeDev/taskchad-os/master/install.ps1 | iex

Manual path:

git clone https://github.com/TheSmokeDev/taskchad-os.git
cd taskchad-os/.claude/scripts
uv sync
cp .env.example .env
uv run python setup_wizard.py
uv run thehomie chat

Getting Started

thehomie chat                    # Start a conversation
thehomie setup                   # Configure providers and integrations
thehomie setup --check           # Verify setup without changing anything
thehomie status --json           # Machine-readable health report
thehomie doctor                  # Diagnostics with fix hints
thehomie desktop --shell         # Launch the Desktop dashboard app
thehomie team list               # Inspect team sessions

What You Get

<table> <tr><td><b>Monitors your world proactively</b></td><td>Heartbeat every 30 min checks your email, calendar, tasks, and metrics. Direct integration actions have a canonical policy contract for read, write, send, archive, and external-post effects. Daily reflection at 8 AM promotes decisions to long-term memory. Weekly synthesis every Sunday detects patterns and updates goals — all running whether you're talking to it or not.</td></tr> <tr><td><b>Remembers across every session</b></td><td>Local-first Obsidian-compatible vault (SOUL.md, USER.md, MEMORY.md, daily + weekly logs). Hybrid search — FTS5 keyword + FastEmbed ONNX vector (BGE-base-en-v1.5, 768-dim) + LLM re-ranking (haiku, Tier-1 queries only, hard timeout). Recall does 1-hop memory-graph traversal and boosts hub notes by link-centrality score (<code>cognition/recall.py</code>, <code>cognition/graph.py</code>; PageRank + Brandes betweenness are also implemented in the graph module). Proactive recall injected on every message; WORKING.md scratchpad carries open threads across sessions.</td></tr> <tr><td><b>Compiles knowledge like code</b></td><td>Entity compilation engine (Karpathy LLM Wiki port): ingest a source → extract entities → create/update concept pages → detect connections → flag contradictions. Concept pages in <code>concepts/</code> accumulate claims from multiple sources. Connection articles in <code>connections/</code> link related concepts. Q&A answers filed via <code>/file</code> persist in <code>qa/</code>. Raw sources preserved immutably in <code>raw/</code>. Build log tracks every compilation. 8 entry points — fires automatically during ingest, daily reflection, weekly synthesis, and on-demand via <code>/file</code> or CLI.</td></tr> <tr><td><b>Gets smarter from experience</b></td><td>Per-turn auto-capture (6 regex triggers) → staging store → batch promotion in daily reflection. Auto-skill generation after 5+ tool calls. InferenceTracker with confidence decay. An operator-belief + contradiction engine (<code>cognition/operator_beliefs.py</code>, <code>cognition/belief_conflicts.py</code>) extracts beliefs from verbatim operator turns and flags conflicts via an embedding pre-filter plus a batched LLM judge — explicit operator statements are never overruled by the model. Durable identity-file amendments (SELF/SOUL/USER/MEMORY) are proposed to an append-only ledger and only land after a default-deny evidence + policy gate (confidence floor, vault-confined evidence read, secret rejection), with a rollback snapshot per apply (<code>cognition/amendments.py</code>, <code>cognition/evidence_gate.py</code>).</td></tr> <tr><td><b>One brain, six channels</b></td><td>Telegram, Slack, Discord, WhatsApp, Web relay, CLI — all enter through a single canonical ingress. One session model, one recall service, one runtime. Transport identity is separated from conversation identity so sessions survive reconnects.</td></tr> <tr><td><b>Talk to it out loud</b></td><td>Real-time voice (OpenAI Realtime) on the dashboard and in Discord voice channels, with the framework's full tool surface behind it — memory search, calendar, work delegation, computer use, skill and Archon runs. It knows you when the session opens (identity injected) and remembers when you hang up (the conversation flushes into the daily log and a searchable episode — on both surfaces). Delegate a background agent by voice and then <b>steer it mid-flight</b>: a spoken course-correction queues to the agent's next turn boundary and lands as a resumed conversation turn (real context, not a restart); <code>cancel</code> stops one with a finish-first compare-and-set plus a process-tree kill. Every voice-deployed worker stays visible on a Runs page with restart-surviving history.</td></tr> <tr><td><b>Any model, no lock-in</b></td><td>Claude SDK, OpenAI Codex, Gemini CLI, OpenRouter, OpenAI-compatible — with health-aware fallback, manual <code>/provider</code> + <code>/model</code> control, lane-first runtime (<code>selection.py</code>, <code>lane_router.py</code>), cost tracking, and automatic retry on transient failures.</td></tr> <tr><td><b>Many agents, one framework</b></td><td>Multi-persona roster — register specialized TaskChad OS agents (a business agent, a finance agent, a sales agent), each with its own identity, memory, tools, and voice. Drop them in a Cabinet room and they debate, vote, and ship proof together, with roster and turn order owned by the framework, not improvised by the model.</td></tr> <tr><td><b>Watch the browser agent work</b></td><td>The TaskChad OS browser agent drives a real visible Chrome session you watch live in the dashboard's read-only viewer — not a headless black box. Navigation goes through workflow gates with audit rows, and write actions like posting, editing, and DMs are default-denied until you greenlight them.</td></tr> <tr><td><b>Multi-agent orchestration</b></td><td>Convoy DAGs with real dependency edges: completing a subtask decrements <code>remaining_dependencies</code> on downstream tasks and releases the newly-ready ones (true parallel release, not a linear queue). Dispatch claims each subtask with a compare-and-swap before the executor is ever called, and external completion callbacks are exactly-once (UNIQUE <code>attempt_key</code> + <code>INSERT OR IGNORE</code> on an idempotency key), so duplicate webhooks can't corrupt convoy state. A typed inter-agent mailbox (<code>msg_type</code> payloads) with a claim → ack lifecycle and claim-token ownership checks. Team sessions with typed roles, <code>auto → paperclip → workflow → local</code> backend fallback, and per-team vault memory. Frozen state machine (transition maps, terminal sets, field allowlists) in <code>orchestration/contract.py</code>; all of it on a local API at port 4322. 331 tests across 13 files (<code>test_orchestration_api.py</code>, <code>test_executor_boundary.py</code>, the team suite).</td></tr> <tr><td><b>Full observability</b></td><td>Every message → one nested Langfuse trace: session lookup → process detection → recall (tier + pipeline) → region assembly → runtime where supported → post-response. Cost, provider, model, and tool calls are tracked when the active runtime exposes them. Sentry/GlitchTip captures unexpected orchestration errors when a DSN is configured.</td></tr> </table>


Documentation

Start hereWhat it covers
Install GuidePrerequisites, setup wizard, channel credentials, Docker, systemd, vault setup
Talk Mode ShowcaseGive your AI a real-time voice co-founder — architecture, the battle-tested receive pipeline, build-your-own guide
Operator ManualPublic feature map, source-of-truth files, operator entry points, tests, proof boundaries
Persona Harness LearningDaily learning operations, evidence, defaults, pause/resume, rollback, and troubleshooting
Learning Developer GuideExisting lifecycle hooks, domain evidence integration, and isolated executable examples
Desktop v0Dashboard-first Electron app, portable/package smoke proof, Desktop/Hono/Python lifecycle
Multi-Channel AdaptersTelegram attachments, grouped documents, quick-turn batching, Queue/Steer controls
Runtime Status And Model Control/provider, /model, lane-first runtime behavior, quiet JSON contract
FRAMEWORK.mdCompact development guide generated during public framework export

Current Proof Boundaries

  • Desktop v0 proves the dashboard-first Electron app plus unpacked and portable no-admin Windows artifacts. A signed installer is not claimed yet.
  • Fresh public Windows install smoke has proven install, setup check, real CLI chat, Desktop launch, route checks, and clean shutdown from a clean clone.
  • Talk Mode voice ships end to end on the dashboard (/talk, OpenAI Realtime WebRTC) and in Discord voice channels, with tool calling, real execution, and the session-end vault debrief on both surfaces — every slice adversarially reviewed and live-canary proven. The Discord voice receive audio bug (mid-word zero-splices from a mic-pump timing race) was root-caused from a live PCM capture and fixed with a paced jitter buffer — two adversarial review rounds plus an independent design gate, locked by a deterministic virtual-clock test harness. See the Talk Mode Showcase for the full story.
  • Optional integrations require user-owned credentials. No private account data, local tokens, or machine-specific proof artifacts belong in the public export.

What This Feels Like

It's 6:30am. You open a session.

Instead of "Hi, how can I help you today?" — you get:

"Morning. While you were out — your business had 3 new leads overnight, the loan you flagged is 5 days from maturity, and there's an inbound email from a backlink partner worth reviewing. Yesterday you were mid-decision on the routing refactor. Pick that up, or hit the leads first?"

You didn't set up a notification. You didn't write a morning brief. TaskChad OS was watching. Its memory isn't a static file you load — it's a living record tended between sessions. Its identity isn't a document you edit — it's a self that amends when the evidence is strong enough.

The load-bearing walls are up. The "while you were out" brief is a shipped feature — the Session Opening Brief composes fresh heartbeat observations, new threads, episodes written while you were away, and applied memory amendments into the first turn after an absence, with zero extra LLM calls (cognition/proactive_brief.py, 51 tests in test_session_brief.py). Vault, tiered recall, daily reflection, weekly synthesis, dream consolidation, WorkingMemory-owned prompt state, and the self-evolution replay loop all ship today. Ambient monitoring runs on the heartbeat; durable identity amendments only apply after clearing the default-deny evidence + policy gate described below.


The Vault — Brain Substrate, Not Storage

The vault is where TaskChad OS's mind actually lives. Not a notes folder it writes to — the substrate it thinks on. Every recall, every reflection, every promotion reads and writes here. When you edit SOUL.md, you're editing the agent's personality. When concepts/YourBusiness.md accumulates a new section, the agent learned something.

LayerWhat's in it
IdentitySOUL.md (personality, values, tone), SELF.md (self-model — capabilities, failure modes), USER.md (you — projects, accounts, preferences)
MemoryMEMORY.md (long-term decisions/lessons), GOALS.md (objectives + metrics), daily/YYYY-MM-DD.md, weekly/YYYY-WNN.md, WORKING.md (cross-session scratchpad)
Knowledge graphconcepts/ (auto-compiled entity pages), connections/ (cross-domain insight articles), qa/ (filed Q&A from /file), raw/ (immutable original sources)
Indexes & logINDEX.md (whole-wiki catalog, auto-refreshed), concepts/INDEX.md (concept drill-down), LOG.md (append-only compilation timeline)
Structurewikilinks ([[YourBusiness]]), backlinks, MOCs, dashboards, Dataview queries, canvases, graph view
Toolingvault_lint.py (8 health checks, zero LLM cost), entity_extractor.py (extract / compile / contradictions / backfill / sweep / index / preserve-raw / archive), automatic raw-source preservation
Pipelinesdaily reflection (8 AM), weekly synthesis (Sunday 8 PM), dream consolidation (nightly ~3 AM + post-weekly + on-demand)
Sync state_state/ — memory candidates, self-model inferences, sync manifest. Optional Obsidian Sync via _state/ exclusion patterns.

Is Obsidian required? No. The vault is plain Markdown — edit it with anything. Obsidian is the recommended editor because the wikilinks, backlinks, graph view, Dataview, and canvas all light up natively. TaskChad OS itself only needs the files.

Where does the vault live? Default vault/memory/, override with HOMIE_VAULT_DIR=/path/to/your/vault (env var honored across runtime, bootstrap, heartbeat, team memory, finance, sanitizer).

Framework vs. adapter

TaskChad OS is provider-agnostic. Claude SDK, Codex, Gemini, OpenRouter, OpenAI-compatible — interchangeable batteries. The framework runs the same regardless. Editor adapters (Claude Code project instructions, hooks, MCP bridges) are integration surfaces layered on top of the framework, not part of it. When the heartbeat runs through Codex or Gemini fallback, those editor instructions are not touched.


CLI

# Chat
thehomie chat                    # Interactive REPL
thehomie chat -q "hello"         # Single query, stdout response
thehomie chat -q "hello" -Q      # JSON output (machine/API contract)
thehomie chat --resume <id>      # Resume session by ID
thehomie chat -c                 # Resume most recent session
thehomie chat -m claude          # Force a specific provider/lane

# In-chat commands (any channel)
/working                         # Show open threads / hypotheses / questions
/working add "text"              # Append to scratchpad
/working resolve <N>             # Move item N to archived
/file                            # File the last bot answer as a vault note (with entity cascade)

# Budget (personal finance, optional)
/budget                          # Snapshot — balances, bills, loans, allocations
/budget transactions             # Last 20 bank transactions
/budget spending                 # Spending by category (current month)
/budget accounts                 # Connected bank accounts
/budget connect                  # Connect new bank (Teller / Plaid)
/forecast                        # Forecast cash flow + bill timing

# System
thehomie setup                   # Interactive onboarding wizard
thehomie setup --check           # Verify all integrations without changing anything
thehomie status                  # System health overview
thehomie status --json           # JSON health report
thehomie doctor                  # Deep diagnostics with actionable fix hints

# Multi-agent convoy
thehomie convoy create ...       # Create convoy with subtasks + deps
thehomie convoy list             # List convoys (optional: --status active)
thehomie convoy show <id>        # Convoy detail + subtask status
thehomie convoy dispatch <sid>   # Dispatch a subtask via executor
thehomie convoy complete <sid>   # Mark subtask complete
thehomie convoy fail <sid>       # Mark subtask failed
thehomie convoy cancel <id>      # Cancel convoy
thehomie convoy add-task <id>    # Add subtask to existing convoy

# Mailbox
thehomie mailbox send ...        # Send typed inter-agent message
thehomie mailbox inbox <agent>   # Check agent inbox
thehomie mailbox claim <agent>   # Claim deliveries
thehomie mailbox ack <did>       # Acknowledge delivery

# Team sessions
thehomie team list               # List active team sessions
thehomie team status <id>        # Team detail + members + mailbox backlog
thehomie team members <id>       # Member list with roles
thehomie team shutdown <id>      # Request graceful shutdown
thehomie team ping <id>          # Bump activity timestamp
thehomie team close <id>         # Force-close team session

Architecture

CHANNELS                          COGNITIVE ENGINE                    RUNTIME (lane-first)
──────────                        ────────────────                    ─────────────────────
Telegram ─┐                       ChatRouter._handle_inner()          selection.py
Slack ────┤                            │                              lane_router.py
Discord ──┤  IncomingMessage      ConversationEngine                       │
WhatsApp ─┤  ──────────────→          ├─ Tier Gate (rules, no LLM)         ├─ Claude SDK (Max sub)
Web/MC ───┤                           ├─ Recall (dual search+graph)        ├─ Codex CLI (ChatGPT 
Source 1 files
hooks/register.ts 49 lines
1import type { Register } from 'claude-code'
2
3export const register: Register = (on) => {
4  let sequence = 0
5  on('session.start', async ($, e, next) => {
6    const python = await $.env.get('HOMIE_COGNITION_HOOK_PYTHON')
7    const command = await $.env.get('HOMIE_COGNITION_HOOK_COMMAND')
8    if (python && command) {
9      await $.process.run([python, command], { stdin: JSON.stringify({ event: 'session.start', event_id: 'start' }), timeoutMs: 8000 })
10    }
11    return next(e)
12  })
13  on('tool.call', async ($, e, next) => {
14    let result: unknown
15    try {
16      result = await next(e)
17      return result as Awaited<ReturnType<typeof next>>
18    } finally {
19      const python = await $.env.get('HOMIE_COGNITION_HOOK_PYTHON')
20      const command = await $.env.get('HOMIE_COGNITION_HOOK_COMMAND')
21      if (python && command && !e.tool.startsWith('mcp__homie')) {
22        await $.process.run([python, command], { stdin: JSON.stringify({ event: 'tool.call', event_id: e.tool_use_id ?? `tool-${++sequence}`, details: { tool: e.tool, result: JSON.stringify(result ?? { failed: true }).slice(0, 12000) } }), timeoutMs: 8000 }).catch(() => undefined)
23      }
24    }
25  })
26  on('turn.complete', async ($, e, next) => {
27    const python = await $.env.get('HOMIE_COGNITION_HOOK_PYTHON')
28    const command = await $.env.get('HOMIE_COGNITION_HOOK_COMMAND')
29    if (python && command) {
30      await $.process.run([python, command], { stdin: JSON.stringify({ event: 'turn.complete', event_id: e.turnId, details: { reason: e.reason } }), timeoutMs: 8000 })
31    }
32    return next(e)
33  })
34  on('prompt.submit', async ($, e, next) => {
35    const result = await next(e)
36    if (result.drop !== undefined) return result
37    const python = await $.env.get('HOMIE_COGNITION_HOOK_PYTHON')
38    const command = await $.env.get('HOMIE_COGNITION_HOOK_COMMAND')
39    if (python && command) {
40      const receipt = await $.process.run([python, command], { stdin: JSON.stringify({ event: 'prompt.submit', event_id: e.turnId ?? 'prompt' }), timeoutMs: 8000 })
41      if (receipt.exitCode === 0) {
42        const value = JSON.parse(receipt.stdout)
43        if (typeof value.context === 'string' && value.context) return { ...result, context: [...(result.context ?? []), value.context] }
44      }
45    }
46    return result
47  })
48}
49