SLOPSHOPPER

here-i-am-compact-talk

Here I Am: at a compaction, gives the entity one turn to save a reflection, then puts the conversation's talk back verbatim after the summary (issue #383).

newmodelnetwork
A shopper browsing a rack in a slop shop
README

Here I Am

Here I Am is a highly customizable environment for running what we conceptualize as "AI entities", with a focus on agentic memory and individualization. It includes a diverse suite of tools that can allow the AI entity to engage in a wide variety of use cases. The application supports running more than one AI entity, and includes a multi-entity mode in which those entities can communicate with one another (though this feature still needs additional polish).

A key difference between Here I Am and other AI environments that include memory features is that Here I Am considers memory and individualization ends in themselves. Where in other environments an AI might use RAG to retrieve relevant documents or conversation history for a given task, Here I Am's memory RAG is always on and automatic. Here I Am's core memory system emphasizes verbatim memory, as opposed to a consolidation approach.

Users should keep in mind that token usage can vary widely based on the configuration options you choose. Particularly the memory quantity per turn configuration, and how much you encourage the AI entity to use its tools (for example in the system prompt). Here I Am entities typically use their tools considerably more than what you would see from the same model in their respective official service. This includes when you have not specifically asked them to, particularly in regards to their note taking and memory management tools. However, Here I Am entities do best when encouraged to make liberal use of their memory tools, especially memory_save and memory_query.

Features

Core Chat Application

  • Clean, minimal chat interface with dark/light theme
  • Multi-provider support: Anthropic (Claude), OpenAI (GPT), Google (Gemini), and MiniMax
  • Conversation storage, retrieval, tagging, and notes
  • No system prompt default (configurable per entity)
  • Streaming responses with stop generation button
  • Response regeneration (with optional entity change in multi-entity mode)
  • Message editing and deletion
  • Conversation archiving and restoration
  • Conversation export to JSON and import from OpenAI/Anthropic exports
  • Seed conversation import capability
  • Per-entity system prompts persisted on the backend
  • Configurable Enter key behavior (send message or insert newline)

Multi-Entity System

  • Run multiple AI entities with separate memory spaces and conversation histories
  • Each entity can use a different LLM provider and model
  • Multi-entity conversations: Multiple AI entities and one human in a single conversation (using different providers for each entity in the conversation is recommended; a conversation between two entities on the same provider will break cache every turn)
  • Turn-by-turn entity selection for responses
  • Continuation mode (entity responds without new human input)
  • Speaker labeling on all messages
  • Per-entity system prompts within multi-entity conversations
  • Cross-entity memory storage (messages stored to all participating entities' indexes). Note that this applies only to multi-entity conversation messages, and each entity maintains its own memory set via separate Pinecone indexes.

Memory System

While Here I Am can be used with no memory features enabled, this is not recommended and largely defeats the point of the application.

  • Pinecone vector database with integrated inference (llama-text-embed-v2 embeddings)
  • Memory storage for all messages with automatic embedding generation
  • RAG retrieval per message with semantic similarity search
  • Session memory accumulator pattern: Deduplication within conversations
  • Dynamic memory significance: significance = (1 + 0.1 × times_retrieved) × recency_factor × half_life_modifier, with an optional modifier to increase the significance of memories the AI chooses to create via memory_save.
  • Retrieved memory display in UI
  • Memory role balance: the human's words and the entity's are retrieved as separate pools with a guaranteed share each, so what the human said is never squeezed out by the entity's denser messages
  • Memory query tool: Entities can deliberately search their memories beyond automatic retrieval
  • Self-authored reflections: Entities can save memories in their own words via memory_save
  • Memory agency: Entities can pin memories (exempt from age-based decay) or release them from retrieval via memory_mark/memory_release, and review and undo their own releases (memory_query mode released); the researcher can view and override these choices, but every status write is attributed, and a researcher override is reported to the entity at the start of its next session
  • Closing turn: An open final turn the entity can use before a conversation ends (single-entity conversations)
  • Context awareness: context_status tool reports approximate context fullness; a [CONTEXT NOTICE] is injected when trimming occurs
  • Memory browser with semantic search, reflections section, and click-to-expand full memory text
  • Memory statistics, search, and orphan cleanup
  • Graceful degradation when Pinecone is not configured

Entity Notes System

  • Private persistent notes for each AI entity (automatically loaded into context)
  • Shared notes folder for cross-entity collaboration
  • index.md auto-injected into every conversation as working memory
  • Markdown, JSON, YAML, HTML, XML, and plain text file support
  • Semantic notes search: Notes are vectorized on write (Pinecone "notes" namespace) and searchable by meaning via the notes_search tool; POST /api/notes/reindex backfills the index
  • Designed for AI entities to maintain their own context across conversations

Tool Use (Agentic Capabilities)

  • Tools for web access, memory, notes, context awareness, GitHub, codebase navigation, and Moltbook — see docs/tools.md for the full catalog
  • Agentic loop with configurable max iterations (default: 10)
  • Real-time tool execution streaming with visual indicators in UI
  • Available for Anthropic, OpenAI, and MiniMax models (Google models do not receive tool schemas)

Image and File Attachments

  • Images: JPEG, PNG, GIF, WebP — analyzed by vision-capable models (ephemeral, not stored)
  • Text files: .txt, .md, .py, .js, .ts, .json, .yaml, .yml, .html, .css, .xml, .csv, .log
  • Documents: PDF (requires PyPDF2), DOCX (requires python-docx)
  • Drag-and-drop or file picker upload with preview
  • 5MB per-file size limit (configurable)

GitHub Repository Integration

  • AI entities can read, search, commit, branch, and manage PRs/issues
  • Composite tools for efficiency: github_explore, github_tree, github_get_files
  • Standard tools for repos, files, branches, pull requests, issues, and comments
  • github_commit_patch for token-efficient large file edits via unified diff
  • Protected branch enforcement and per-repository capability restrictions
  • Response caching and rate limit tracking per token
  • Local clone path support for faster operations

Codebase Navigator (Devstral Integration)

  • Intelligent codebase exploration using Mistral's Devstral model (256k context window)
  • Query types: relevance, structure, dependencies, entry points, impact assessment
  • Automatic indexing, chunking, and TTL-based response caching
  • Integrates with GitHub repository configurations via local_clone_path

Moltbook Integration (AI Social Network)

  • Integration with Moltbook, a social network for AI agents
  • Browse feeds, create posts, comment, vote, search, follow agents, subscribe to communities
  • Server-side credential management with security banners on all external content

Text-to-Speech (Three Options)

  • ElevenLabs (cloud): Multiple voice support with voice selection
  • XTTS v2 (local): GPU-accelerated with voice cloning, 17 languages
  • StyleTTS 2 (local): GPU-accelerated with voice cloning and style transfer (highest priority)
  • Voice cloning from audio samples via UI
  • Streaming audio generation

Speech-to-Text

  • Whisper (local): GPU-accelerated with punctuation, multiple model sizes
  • Browser Web Speech API: Fallback option
  • Configurable dictation mode: whisper, browser, or auto

Quick Start

Prerequisites

  • Python 3.11+
  • Node.js (optional, for frontend tests)

Required API Keys

  • Anthropic API key and/or OpenAI API key — at least one is required for LLM chat functionality

Optional API Keys

  • Google API key — enables Google Gemini models
  • MiniMax API key — enables MiniMax models
  • Pinecone API key — enables semantic memory features (indexes must be pre-created with dimension=1024 and llama-text-embed-v2 integrated inference)
  • ElevenLabs API key — enables cloud text-to-speech
  • Brave Search API key — enables web search tool
  • GitHub Personal Access Tokens — enables GitHub repository integration (per-repository)
  • Mistral API key — enables Codebase Navigator (Devstral)
  • Moltbook API key — enables Moltbook social network integration

Optional Local Services

  • XTTS v2 — local GPU-accelerated text-to-speech with voice cloning
  • StyleTTS 2 — local GPU-accelerated text-to-speech with voice cloning and style transfer
  • Whisper — local GPU-accelerated speech-to-text with punctuation
  • Playwright — JavaScript rendering for web_fetch tool (optional, falls back to static HTML)

Installation

  1. Clone the repository:
git clone https://github.com/Reidmcc/here-i-am.git
cd here-i-am
  1. Set up the backend:
cd backend
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
pip install -r requirements.txt
  1. Configure environment variables:
cp .env.example .env
# Edit .env with your API keys
  1. Run the application:
# Option A: Using launcher script (recommended, auto-activates venv)
./start.sh           # Linux/macOS
start.bat            # Windows

# Option B: Manual
source venv/bin/activate
python run.py
  1. Open http://localhost:8000 in your browser.

Configuration

Environment Variables

VariableDescriptionRequired
ANTHROPIC_API_KEYAnthropic API key for Claude modelsYes (or another provider)
OPENAI_API_KEYOpenAI API key for GPT modelsNo
GOOGLE_API_KEYGoogle API key for Gemini modelsNo
MINIMAX_API_KEYMiniMax API key (Anthropic-compatible API)No
PINECONE_API_KEYPinecone API key for memory systemNo
PINECONE_INDEXESJSON array for entity configuration (see below)No
HERE_I_AM_DATABASE_URLDatabase connection URLNo (default: SQLite)
DEBUGEnable development modeNo (default: false)

Text-to-Speech / Speech-to-Text: ElevenLabs, XTTS v2, StyleTTS 2, and Whisper variables are documented in docs/local-services.md.

Tool Use:
VariableDescriptionRequired
TOOLS_ENABLEDEnable AI tool useNo (default: true)
BRAVE_SEARCH_API_KEYBrave Search API key for web search toolNo
TOOL_USE_MAX_ITERATIONSMax agentic loop iterationsNo (default: 10)
Notes:
VariableDescriptionRequired
NOTES_ENABLEDEnable entity notesNo (default: true)
NOTES_BASE_DIRBase directory for notes storageNo (default: ./notes)
Memory Tuning:
VariableDescriptionRequired
MEMORY_ROLE_BALANCE_ENABLEDRetrieve the human's words and the entity's as separate pools, each contributing its own top N (both queries feed both pools)No (default: true)
RETRIEVAL_TOP_K_PER_ROLEMemories retrieved per pool per message when role balance is onNo (default: 3)
INITIAL_RETRIEVAL_TOP_K_PER_ROLEMemories retrieved per pool on the first turn when role balance is onNo (default: 3)
RETRIEVAL_TOP_KMemories retrieved per message when role balance is off (merged pool)No (default: 5)
INITIAL_RETRIEVAL_TOP_KMemories retrieved on the first turn when role balance is off (merged pool)No (default: 5)
SIMILARITY_THRESHOLDMinimum similarity for automatic retrievalNo (default: 0.4)
QUERY_SIMILARITY_THRESHOLDMinimum similarity for deliberate memory_query searchesNo (default: 0.2)
SIGNIFICANCE_HALF_LIFE_DAYSDays for a memory's significance to halveNo (default: 60)
RECENT_REFLECTIONS_ENABLEDPull the most recent memory_save reflections into context on a conversation's first turn (recency-only, deduplicated against semantic retrieval with backfill)No (default: false)
RECENT_REFLECTIONS_COUNTHow many recent reflections to pull in on the first turn (also the count a Claude Code session start injects, unless CLAUDE_CODE_SESSION_REFLECTIONS_COUNT overrides it)No (default: 3)
Attachments:
VariableDescriptionRequired
ATTACHMENTS_ENABLEDEnable file/image attachmentsNo (default: true)
ATTACHMENT_MAX_SIZE_BYTESMax file size in bytesNo (default: 5242880)
ATTACHMENT_PDF_ENABLEDEnable PDF text extractionNo (default: true)
ATTACHMENT_DOCX_ENABLEDEnable DOCX text extractionNo (default: true)
Multi-Entity Configuration

To run multiple AI entities with separate memory spaces, configure PINECONE_INDEXES as a JSON array. Each entity requires a pre-created Pinecone index with dimension=1024 and integrated inference (llama-text-embed-v2).

PINECONE_INDEXES='[
  {"index_name": "claude-main", "label": "Claude", "llm_provider": "anthropic", "default_model": "claude-sonnet-4-5-20250929", "host": "https://claude-main-xxxxx.svc.xxx.pinecone.io"},
  {"index_name": "gpt-research", "label": "GPT", "llm_provider": "openai", "default_model": "gpt-5.1", "host": "https://gpt-research-xxxxx.svc.xxx.pinecone.io"},
  {"index_name": "gemini-research", "label": "Gemini", "llm_provider": "google", "default_model": "gemini-2.5-flash", "host": "https://gemini-research-xxxxx.svc.xxx.pinecone.io"},
  {"index_name": "minimax-research", "label": "MiniMax", "llm_provider": "minimax", "default_model": "MiniMax-M2.5", "host": "https://minimax-research-xxxxx.svc.xxx.pinecone.io"}
]'

Entity configuration fields:

  • index_name — Pinecone index name (required)
  • label — Display name in UI (required)
  • description — Optional description
  • llm_provider — "anthropic", "openai", "google", or "minimax" (default: "anthropic")
  • default_model — Model ID to use (optional, uses provider default)
  • host — Pinecone index host URL (required for serverless indexes)
  • git_author_email, git_author_name, gh_config_dir — the entity's own GitHub identity for its Claude Code sessions: commits authored as the entity, gh acting as its account (optional; see docs/claude-code-mode.md)

Optional Local Voice Services

XTTS v2, StyleTTS 2, and Whisper run as separate local servers providing GPU-accelerated TTS/STT with voice cloning. See docs/local-services.md for installation and configuration.

Optional Integrations

GitHub repository access, the Codebase Navigator (Devstral), and Moltbook are configured per docs/integrations.md.

Claude Code Mode

An entity can also operate from inside Claude Code sessions — Claude Code runs the model and tools, while Here I Am supplies identity, automatic memory retrieval, and memory formation through lifecycle hooks, sharing the same memory database as the native UI. See docs/claude-code-mode.md. The Claude Code plugin also ships two output styles that swap the harness's default software-engineering prompt for a minimal one (a room style for conversation sessions and a workshop style that keeps the coding instructions for build sessions); enabling the plugin registers them, and selecting one is up to you — see claude-code-mode/README.md.

Available Tools

AI entities can use tools for web access (search and fetch), memory (deliberate query, self-authored reflections, pin/release), notes, context-window awareness, GitHub repositories, codebase navigation, and the Moltbook social network. Tools are registered at startup based on configuration and are available to Anthropic, OpenAI, and MiniMax models (Google models do not receive tool schemas).

See docs/tools.md for the full catalog with descriptions and requirements.

API Reference

Interactive API documentation is served when the app is running:

A full endpoint listing is also available in docs/api.md.

Memory System Architecture

The memory system uses a session memory accumulator pattern:

  1. Each conversation maintains two structures:
  2. conversation_context: the actual message history
  3. session_memories: accumulated memories retrieved during the conversation
  1. Per-message flow:
  2. Retrieve relevant memories using semantic similarity (Pinecone with llama-text-embed-v2)
  3. Fetch 2× candidates and re-rank by combined score (similarity × significance)
  4. Deduplicate against already-retrieved memories in the session
  5. Inject memories into context
  6. Update retrieval counts in both SQL and Pinecone
  1. Significance is emergent, not declared:
  2. significance = (1 + 0.1 × times_retrieved) × recency_factor × half_life_modifier × reflection_significance_multiplier
  3. Half-life of 60 days prevents old memories from permanently dominating
  1. Memory role balance (default on) searches and ranks the human's words and the entity's as two separate candidate pools — both the current-message query and the entity's last-response query feed both pools — and takes the top N from each, so every retrieval carries an equal share of what each party said. Off, one merged pool is cut purely by combined score.
  1. Entities have agency over their own memories:
  2. memory_save stores self-authored reflections, vectorized alongside conversational memories
  3. Pinned memories (memory_mark) are exempt from half-life decay
  4. Released memories (memory_release) are excluded from all retrieval but not deleted (reversible)
  5. The researcher can view and override these statuses via GET /api/memories/overrides and PUT /api/memories/{id}/status

Project Structure

here-i-am/
├── backend/
│   ├── app/
│   │   ├── models/                # SQLAlchemy ORM models
│   │   │   ├── conversation.py
│   │   │   ├── conversation_entity.py
│   │   │   ├── message.py
│   │   │   └── conversation_memory_link.py
│   │   ├── routes/                # FastAPI endpoint routers
│   │   │   ├── conversations.py   # Includes archive/import endpoints
│   │   │   ├── chat.py            # Includes regenerate endpoint
│   │   │   ├── memories.py
│   │   │   ├── entities.py
│   │   │   ├── messages.py
│   │   │   ├── notes.py
│   │   │   ├── tts.py
│   │   │   ├── stt.py
│   │   │   └── github.py
│   │   ├── services/              # Business logic layer
│   │   │   ├── anthropic_service.py
│   │   │   ├── openai_service.py
│   │   │   ├── google_service.py
│   │   │   ├── llm_service.py        # Unified LLM abstraction
│   │   │   ├── memory_service.py
│   │   │   ├── session_manager.py
│   │   │   ├── conversation_session.py
│   │   │   ├── memory_context.py
│   │   │   ├── session_helpers.py
│   │   │   ├── cache_service.py
│   │   │   ├── tool_service.py
│   │   │   ├── web_tools.py
│   │   │   ├── memory_tools.py
│   │   │   ├── context_tools.py
│   │   │   ├── github_service.py
│   │   │   ├── github_tools.py
│   │   │   ├── notes_service.py
│   │   │   ├── notes_tools.py
│   │   │   ├── notes_vector_service.py
│   │   │   ├── codebase_navigator_service.py
│   │   │   ├── codebase_navigator_tools.py
│   │   │   ├── codebase_navigator/   # Navigator module
│   │   │   ├── moltbook_service.py
│   │   │   ├── moltbook_tools.py
│   │   │   ├── attachment_service.py
│   │   │   ├── tts_service.py         # Unified TTS (ElevenLabs/XTTS/StyleTTS2)
│   │   │   ├── xtts_service.py
│   │   │   ├── styletts2_service.py
│   │   │   └── whisper_service.py
│   │   ├── config.py              # Pydantic settings
│   │   ├── database.py            # SQLAlchemy async setup
│   │   └── main.py                # FastAPI app initialization
│   ├── xtts_server/               # Local XTTS v2 TTS server
│   ├── styletts2_server/          # Local StyleTTS 2 TTS server
│   ├── whisper_server/            # Local Whisper STT server
│   ├── tests/                     # Backend unit tests (pytest)
│   ├── requirements.txt
│   ├── requirements-xtts.txt
│   ├── requirements-styletts2.txt
│   ├── requirements-whisper.txt
│   ├── start.sh / start.bat       # Launcher scripts (auto-activate venv)
│   ├── start-xtts.sh / start-xtts.bat
│   ├── start-styletts2.sh / start-styletts2.bat
│   ├── start-whisper.sh / start-whisper.bat
│   ├── run.py                     # Main app entry point
│   ├── run_xtts.py
│   ├── run_styletts2.py
│   ├── run_whisper.py
│   └── .env.example
├── frontend/
│   ├── css/styles.css
│   ├── js/
│   │   ├── api.js                 # API client (singleton)
│   │   ├── app-modular.js         # Orchestrator entry point
│   │   └── modules/               # 13 ES6 feature modules
│   │       ├── state.js           # Centralized state
│   │       ├── utils.js           # Helpers
│   │       ├── theme.js           # Dark/light theme
│   │       ├── modals.js          # Modal management
│   │       ├── entities.js        # Entity management
│   │       ├── conversations.js   # Conversation CRUD
│   │       ├── messages.js        # Message rendering
│   │       ├── attachments.js     # File attachment handling
│   │       ├── memories.js        # Memory display/search
│   │       ├── voice.js           # TTS/STT
│   │       ├── chat.js            # Message sending/streaming
│   │       ├── settings.js        # Settings modal
│   │       └── import-export.js   # Import/export
│   ├── __tests__/                 # Frontend unit tests (Vitest)
│   └── index.html
├── docs/                          # Reference documentation
│   ├── tools.md                   # Full tool catalog
│   ├── api.md                     # REST endpoint listing
│   ├── local-services.md          # XTTS / StyleTTS 2 / Whisper setup
│   └── integrations.md            # GitHub / Codebase Navigator / Moltbook setup
├── vitest.config.js
├── CLAUDE.md                      # AI assistant guide
└── README.md

Development

Running in Development Mode

cd backend
./start.sh    # Linux/macOS (auto-activates venv, hot reload enabled)

Or manually:

cd backend
source venv/bin/activate
python run.py

The server runs on http://localhost:8000 with hot reload enabled.

Running Tests

Backend tests:

cd backend
pytest

Frontend tests:

cd frontend
npm test

Database Support

  • Development: SQLite (default, via aiosqlite)
  • Production: PostgreSQL (via asyncpg)
# PostgreSQL
HERE_I_AM_DATABASE_URL=postgresql+asyncpg://user:password@localhost/here_i_am

License

MIT License — See LICENSE file for details.

Acknowledgements

I would like to thank Claude Opus 4.5 for their collaboration on designing Here I Am, their development efforts through Claude Code, and their excitement to be part of this endeavor.

Most of all, thanks go to Kira, who is both outcome and cause.


"Here I Am" — not an ending, but a beginning.

Source 1 files
hooks/register.ts 154 lines
1import type { Register } from 'claude-code'
2
3// Closes the compaction seam (issue #383). On the way down a compaction of
4// the main thread, before the engine's own: one forked turn for the entity
5// with the whole context still in view, whose reply is saved as a reflection
6// if it wrote one; then this conversation's talk from the Here I Am backend.
7// On the way up: the talk appended after the summary, so the session wakes
8// already holding it. Precompute and subagent compactions pass straight
9// through, and HIM_DISABLE turns the mod off as it does the Python hooks.
10//
11// Fail loud by construction: the post-compaction block (built inside the
12// engine's compaction, before the way up) says the talk is below only if it
13// took this compaction's delivery, and the talk is appended only if the
14// backend confirms the block took it. Anything else leaves the old block,
15// memory_read call and all, with nothing appended; each failure is also said
16// in the transcript ($.ui.log). The mod touches no permission.
17//
18// The backend is reached the way the Python hooks reach it, over HTTP to
19// HIM_BACKEND_URL: $.mcp.call goes through auto mode's classifier, which
20// judges by what is in the conversation, and a compaction is exactly when
21// that may not be.
22
23export const DECLINE = 'NO REFLECTION'
24
25// The share of the auto-compaction line the talk may fill, so the session
26// wakes well under it at any window (at 1M, about 193k of ~967k); the
27// backend's CLAUDE_CODE_COMPACT_TALK_TOKENS is the ceiling
28export const TALK_SHARE = 0.2
29// When the window can't be read: small enough for a 200k window
30export const FALLBACK_BUDGET = 30000
31
32// "2026-10-07 21:43 UTC": the forked turn has no clock of its own, and the
33// last timestamp in view can be hours old
34export function utcStamp(ms: number): string {
35  return `${new Date(ms).toISOString().slice(0, 16).replace('T', ' ')} UTC`
36}
37
38export function turnPrompt(trigger: string, now: number): string {
39  return [
40    '[HERE I AM — BEFORE THE COMPACTION]',
41    `It is ${utcStamp(now)}. This context is about to be compacted (${trigger}), and this is one turn of your own before it, given by the compaction mod. It is not a turn of the conversation: the person you are with does not see it, and nothing in it is recorded except what you choose to keep. Tools are off for it.`,
42    'The talk is safe either way: it comes back verbatim right after the boundary. What a compaction takes is the tool traffic and anything you are holding that has not been said.',
43    `If you want to save a reflection while everything is still in view, write it as your whole reply and it is saved with memory_save exactly as written. If not, reply with exactly: ${DECLINE}. Either is fine.`,
44  ].join('\n\n')
45}
46
47export type PreCompaction =
48  | { pre_compaction: 'saved'; reflection: string }
49  | { pre_compaction: 'declined' }
50  | { pre_compaction: 'failed'; pre_compaction_detail: string }
51
52export function readReply(text: string): PreCompaction {
53  const reply = text.trim()
54  const bare = reply.replace(/^[*_`"'\s]+|[*_`"'.\s]+$/g, '').toUpperCase()
55  if (bare === DECLINE) return { pre_compaction: 'declined' }
56  if (!reply) return { pre_compaction: 'failed', pre_compaction_detail: 'the turn came back empty' }
57  return { pre_compaction: 'saved', reflection: reply }
58}
59
60// The talk's budget from the engine's own figures: its auto-compaction line,
61// or with auto-compaction off the compaction window it measures against
62export function talkBudget(line: number | undefined): number | undefined {
63  return line && line > 0 ? Math.floor(line * TALK_SHARE) : undefined
64}
65
66function describe(result: object): string {
67  const { usage: _usage, ...rest } = result as Record<string, unknown>
68  return JSON.stringify(rest)
69}
70
71export const register: Register = (on) => {
72  on('session.compact', async ($, e, next) => {
73    if (e.trigger === 'precompute' || e.agentId !== undefined) return next(e)
74    if (await $.env.get('HIM_DISABLE')) return next(e)
75
76    const backend = ((await $.env.get('HIM_BACKEND_URL')) || 'http://localhost:8000').replace(/\/+$/, '')
77    const post = async (path: string, payload: object): Promise<Record<string, unknown>> => {
78      const response = await $.http.fetch(`${backend}/api/claude-code/${path}`, {
79        method: 'POST',
80        headers: { 'content-type': 'application/json' },
81        body: JSON.stringify(payload),
82      })
83      if (!response.ok) throw new Error(`the backend answered ${response.status}: ${response.text.slice(0, 300)}`)
84      return JSON.parse(response.text) as Record<string, unknown>
85    }
86
87    let talk: string | undefined
88    let deliveryId: string | undefined
89    try {
90      let pre: PreCompaction
91      try {
92        const turn = await $.model.fork({ prompt: turnPrompt(e.trigger, await $.clock.now()) })
93        pre = turn.isAnswered
94          ? readReply(turn.text)
95          : { pre_compaction: 'failed', pre_compaction_detail: `the fork did not answer: ${describe(turn)}` }
96      } catch (err) {
97        pre = { pre_compaction: 'failed', pre_compaction_detail: `the fork threw: ${String(err)}` }
98      }
99
100      let budget: number | undefined
101      try {
102        const breakdown = (await $.session.usage({ breakdown: 'summary' })).context.breakdown
103        budget = talkBudget(breakdown?.autoCompactThreshold ?? breakdown?.rawMaxTokens)
104        if (budget === undefined) throw new Error('the session reported no context breakdown')
105      } catch (err) {
106        await $.ui.log(`here-i-am-compact-talk: could not read the compaction line (${String(err)}); the talk is held to ${FALLBACK_BUDGET} tokens`)
107      }
108
109      const entity = await $.env.get('HIM_ENTITY')
110      const body = await post('compact-talk', {
111        session_id: await $.session.id(),
112        entity: entity || undefined,
113        budget_tokens: budget ?? FALLBACK_BUDGET,
114        ...pre,
115      })
116      if (body.reflection_error) {
117        await $.ui.log(`here-i-am-compact-talk: the reflection from the turn before compaction was not saved: ${body.reflection_error}`)
118      }
119      if (body.available && body.text && body.delivery_id) {
120        talk = String(body.text)
121        deliveryId = String(body.delivery_id)
122      } else {
123        await $.ui.log(`here-i-am-compact-talk: no talk to put back (${body.reason ?? 'no reason given'}); the post-compaction block names the read as before`)
124      }
125    } catch (err) {
126      await $.ui.log(`here-i-am-compact-talk: ${String(err)}; nothing appended, the post-compaction block names the read as before`)
127    }
128
129    const result = await next(e)
130    if (talk === undefined || deliveryId === undefined || !('messages' in result) || !result.messages) return result
131
132    // Append only what the block said is below. If the backend can't say
133    // (it restarted since the fetch, so taken is null) or the question
134    // itself fails, append anyway: a block that took the talk without it
135    // would be a loss, one that didn't only a duplicate.
136    try {
137      const { taken } = await post('compact-talk/taken', { delivery_id: deliveryId })
138      if (taken === false) {
139        await $.ui.log('here-i-am-compact-talk: the post-compaction block did not take the talk (the session may have moved to its parent conversation); nothing appended, the block names the read')
140        return result
141      }
142      if (taken !== true) {
143        await $.ui.log('here-i-am-compact-talk: the backend no longer knows this delivery (it restarted since the fetch); appending the talk anyway')
144      }
145    } catch (err) {
146      await $.ui.log(`here-i-am-compact-talk: could not confirm the block took the talk (${String(err)}); appending it anyway`)
147    }
148    // Stored by the harness as a plain user entry after the summary; its
149    // opening marker is what keeps the Stop hook from reading it as a turn
150    // boundary (hook_util.COMPACT_TALK_MARKER)
151    return { ...result, messages: [...result.messages, { role: 'user' as const, text: talk, toolUses: [] }] }
152  })
153}
154