Here I Am: at a compaction, gives the entity one turn to save a reflection, then puts the conversation's talk back verbatim after the summary (issue #383).

Here I Am is a highly customizable environment for running what we conceptualize as "AI entities", with a focus on agentic memory and individualization. It includes a diverse suite of tools that can allow the AI entity to engage in a wide variety of use cases. The application supports running more than one AI entity, and includes a multi-entity mode in which those entities can communicate with one another (though this feature still needs additional polish).
A key difference between Here I Am and other AI environments that include memory features is that Here I Am considers memory and individualization ends in themselves. Where in other environments an AI might use RAG to retrieve relevant documents or conversation history for a given task, Here I Am's memory RAG is always on and automatic. Here I Am's core memory system emphasizes verbatim memory, as opposed to a consolidation approach.
Users should keep in mind that token usage can vary widely based on the configuration options you choose. Particularly the memory quantity per turn configuration, and how much you encourage the AI entity to use its tools (for example in the system prompt). Here I Am entities typically use their tools considerably more than what you would see from the same model in their respective official service. This includes when you have not specifically asked them to, particularly in regards to their note taking and memory management tools. However, Here I Am entities do best when encouraged to make liberal use of their memory tools, especially memory_save and memory_query.
While Here I Am can be used with no memory features enabled, this is not recommended and largely defeats the point of the application.
significance = (1 + 0.1 × times_retrieved) × recency_factor × half_life_modifier, with an optional modifier to increase the significance of memories the AI chooses to create via memory_save.memory_savememory_mark/memory_release, and review and undo their own releases (memory_query mode released); the researcher can view and override these choices, but every status write is attributed, and a researcher override is reported to the entity at the start of its next sessioncontext_status tool reports approximate context fullness; a [CONTEXT NOTICE] is injected when trimming occursindex.md auto-injected into every conversation as working memory"notes" namespace) and searchable by meaning via the notes_search tool; POST /api/notes/reindex backfills the indexgithub_explore, github_tree, github_get_filesgithub_commit_patch for token-efficient large file edits via unified difflocal_clone_pathwhisper, browser, or autogit clone https://github.com/Reidmcc/here-i-am.git
cd here-i-am
cd backend
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env
# Edit .env with your API keys
# Option A: Using launcher script (recommended, auto-activates venv)
./start.sh # Linux/macOS
start.bat # Windows
# Option B: Manual
source venv/bin/activate
python run.py
| Variable | Description | Required |
|---|---|---|
ANTHROPIC_API_KEY | Anthropic API key for Claude models | Yes (or another provider) |
OPENAI_API_KEY | OpenAI API key for GPT models | No |
GOOGLE_API_KEY | Google API key for Gemini models | No |
MINIMAX_API_KEY | MiniMax API key (Anthropic-compatible API) | No |
PINECONE_API_KEY | Pinecone API key for memory system | No |
PINECONE_INDEXES | JSON array for entity configuration (see below) | No |
HERE_I_AM_DATABASE_URL | Database connection URL | No (default: SQLite) |
DEBUG | Enable development mode | No (default: false) |
Text-to-Speech / Speech-to-Text: ElevenLabs, XTTS v2, StyleTTS 2, and Whisper variables are documented in docs/local-services.md.
| Variable | Description | Required |
|---|---|---|
TOOLS_ENABLED | Enable AI tool use | No (default: true) |
BRAVE_SEARCH_API_KEY | Brave Search API key for web search tool | No |
TOOL_USE_MAX_ITERATIONS | Max agentic loop iterations | No (default: 10) |
| Variable | Description | Required |
|---|---|---|
NOTES_ENABLED | Enable entity notes | No (default: true) |
NOTES_BASE_DIR | Base directory for notes storage | No (default: ./notes) |
| Variable | Description | Required |
|---|---|---|
MEMORY_ROLE_BALANCE_ENABLED | Retrieve the human's words and the entity's as separate pools, each contributing its own top N (both queries feed both pools) | No (default: true) |
RETRIEVAL_TOP_K_PER_ROLE | Memories retrieved per pool per message when role balance is on | No (default: 3) |
INITIAL_RETRIEVAL_TOP_K_PER_ROLE | Memories retrieved per pool on the first turn when role balance is on | No (default: 3) |
RETRIEVAL_TOP_K | Memories retrieved per message when role balance is off (merged pool) | No (default: 5) |
INITIAL_RETRIEVAL_TOP_K | Memories retrieved on the first turn when role balance is off (merged pool) | No (default: 5) |
SIMILARITY_THRESHOLD | Minimum similarity for automatic retrieval | No (default: 0.4) |
QUERY_SIMILARITY_THRESHOLD | Minimum similarity for deliberate memory_query searches | No (default: 0.2) |
SIGNIFICANCE_HALF_LIFE_DAYS | Days for a memory's significance to halve | No (default: 60) |
RECENT_REFLECTIONS_ENABLED | Pull the most recent memory_save reflections into context on a conversation's first turn (recency-only, deduplicated against semantic retrieval with backfill) | No (default: false) |
RECENT_REFLECTIONS_COUNT | How many recent reflections to pull in on the first turn (also the count a Claude Code session start injects, unless CLAUDE_CODE_SESSION_REFLECTIONS_COUNT overrides it) | No (default: 3) |
| Variable | Description | Required |
|---|---|---|
ATTACHMENTS_ENABLED | Enable file/image attachments | No (default: true) |
ATTACHMENT_MAX_SIZE_BYTES | Max file size in bytes | No (default: 5242880) |
ATTACHMENT_PDF_ENABLED | Enable PDF text extraction | No (default: true) |
ATTACHMENT_DOCX_ENABLED | Enable DOCX text extraction | No (default: true) |
To run multiple AI entities with separate memory spaces, configure PINECONE_INDEXES as a JSON array. Each entity requires a pre-created Pinecone index with dimension=1024 and integrated inference (llama-text-embed-v2).
PINECONE_INDEXES='[
{"index_name": "claude-main", "label": "Claude", "llm_provider": "anthropic", "default_model": "claude-sonnet-4-5-20250929", "host": "https://claude-main-xxxxx.svc.xxx.pinecone.io"},
{"index_name": "gpt-research", "label": "GPT", "llm_provider": "openai", "default_model": "gpt-5.1", "host": "https://gpt-research-xxxxx.svc.xxx.pinecone.io"},
{"index_name": "gemini-research", "label": "Gemini", "llm_provider": "google", "default_model": "gemini-2.5-flash", "host": "https://gemini-research-xxxxx.svc.xxx.pinecone.io"},
{"index_name": "minimax-research", "label": "MiniMax", "llm_provider": "minimax", "default_model": "MiniMax-M2.5", "host": "https://minimax-research-xxxxx.svc.xxx.pinecone.io"}
]'
Entity configuration fields:
index_name — Pinecone index name (required)label — Display name in UI (required)description — Optional descriptionllm_provider — "anthropic", "openai", "google", or "minimax" (default: "anthropic")default_model — Model ID to use (optional, uses provider default)host — Pinecone index host URL (required for serverless indexes)git_author_email, git_author_name, gh_config_dir — the entity's own GitHub identity for its Claude Code sessions: commits authored as the entity, gh acting as its account (optional; see docs/claude-code-mode.md)XTTS v2, StyleTTS 2, and Whisper run as separate local servers providing GPU-accelerated TTS/STT with voice cloning. See docs/local-services.md for installation and configuration.
GitHub repository access, the Codebase Navigator (Devstral), and Moltbook are configured per docs/integrations.md.
An entity can also operate from inside Claude Code sessions — Claude Code runs the model and tools, while Here I Am supplies identity, automatic memory retrieval, and memory formation through lifecycle hooks, sharing the same memory database as the native UI. See docs/claude-code-mode.md. The Claude Code plugin also ships two output styles that swap the harness's default software-engineering prompt for a minimal one (a room style for conversation sessions and a workshop style that keeps the coding instructions for build sessions); enabling the plugin registers them, and selecting one is up to you — see claude-code-mode/README.md.
AI entities can use tools for web access (search and fetch), memory (deliberate query, self-authored reflections, pin/release), notes, context-window awareness, GitHub repositories, codebase navigation, and the Moltbook social network. Tools are registered at startup based on configuration and are available to Anthropic, OpenAI, and MiniMax models (Google models do not receive tool schemas).
See docs/tools.md for the full catalog with descriptions and requirements.
Interactive API documentation is served when the app is running:
A full endpoint listing is also available in docs/api.md.
The memory system uses a session memory accumulator pattern:
conversation_context: the actual message historysession_memories: accumulated memories retrieved during the conversationsignificance = (1 + 0.1 × times_retrieved) × recency_factor × half_life_modifier × reflection_significance_multiplier memory_save stores self-authored reflections, vectorized alongside conversational memoriesmemory_mark) are exempt from half-life decaymemory_release) are excluded from all retrieval but not deleted (reversible)GET /api/memories/overrides and PUT /api/memories/{id}/statushere-i-am/
├── backend/
│ ├── app/
│ │ ├── models/ # SQLAlchemy ORM models
│ │ │ ├── conversation.py
│ │ │ ├── conversation_entity.py
│ │ │ ├── message.py
│ │ │ └── conversation_memory_link.py
│ │ ├── routes/ # FastAPI endpoint routers
│ │ │ ├── conversations.py # Includes archive/import endpoints
│ │ │ ├── chat.py # Includes regenerate endpoint
│ │ │ ├── memories.py
│ │ │ ├── entities.py
│ │ │ ├── messages.py
│ │ │ ├── notes.py
│ │ │ ├── tts.py
│ │ │ ├── stt.py
│ │ │ └── github.py
│ │ ├── services/ # Business logic layer
│ │ │ ├── anthropic_service.py
│ │ │ ├── openai_service.py
│ │ │ ├── google_service.py
│ │ │ ├── llm_service.py # Unified LLM abstraction
│ │ │ ├── memory_service.py
│ │ │ ├── session_manager.py
│ │ │ ├── conversation_session.py
│ │ │ ├── memory_context.py
│ │ │ ├── session_helpers.py
│ │ │ ├── cache_service.py
│ │ │ ├── tool_service.py
│ │ │ ├── web_tools.py
│ │ │ ├── memory_tools.py
│ │ │ ├── context_tools.py
│ │ │ ├── github_service.py
│ │ │ ├── github_tools.py
│ │ │ ├── notes_service.py
│ │ │ ├── notes_tools.py
│ │ │ ├── notes_vector_service.py
│ │ │ ├── codebase_navigator_service.py
│ │ │ ├── codebase_navigator_tools.py
│ │ │ ├── codebase_navigator/ # Navigator module
│ │ │ ├── moltbook_service.py
│ │ │ ├── moltbook_tools.py
│ │ │ ├── attachment_service.py
│ │ │ ├── tts_service.py # Unified TTS (ElevenLabs/XTTS/StyleTTS2)
│ │ │ ├── xtts_service.py
│ │ │ ├── styletts2_service.py
│ │ │ └── whisper_service.py
│ │ ├── config.py # Pydantic settings
│ │ ├── database.py # SQLAlchemy async setup
│ │ └── main.py # FastAPI app initialization
│ ├── xtts_server/ # Local XTTS v2 TTS server
│ ├── styletts2_server/ # Local StyleTTS 2 TTS server
│ ├── whisper_server/ # Local Whisper STT server
│ ├── tests/ # Backend unit tests (pytest)
│ ├── requirements.txt
│ ├── requirements-xtts.txt
│ ├── requirements-styletts2.txt
│ ├── requirements-whisper.txt
│ ├── start.sh / start.bat # Launcher scripts (auto-activate venv)
│ ├── start-xtts.sh / start-xtts.bat
│ ├── start-styletts2.sh / start-styletts2.bat
│ ├── start-whisper.sh / start-whisper.bat
│ ├── run.py # Main app entry point
│ ├── run_xtts.py
│ ├── run_styletts2.py
│ ├── run_whisper.py
│ └── .env.example
├── frontend/
│ ├── css/styles.css
│ ├── js/
│ │ ├── api.js # API client (singleton)
│ │ ├── app-modular.js # Orchestrator entry point
│ │ └── modules/ # 13 ES6 feature modules
│ │ ├── state.js # Centralized state
│ │ ├── utils.js # Helpers
│ │ ├── theme.js # Dark/light theme
│ │ ├── modals.js # Modal management
│ │ ├── entities.js # Entity management
│ │ ├── conversations.js # Conversation CRUD
│ │ ├── messages.js # Message rendering
│ │ ├── attachments.js # File attachment handling
│ │ ├── memories.js # Memory display/search
│ │ ├── voice.js # TTS/STT
│ │ ├── chat.js # Message sending/streaming
│ │ ├── settings.js # Settings modal
│ │ └── import-export.js # Import/export
│ ├── __tests__/ # Frontend unit tests (Vitest)
│ └── index.html
├── docs/ # Reference documentation
│ ├── tools.md # Full tool catalog
│ ├── api.md # REST endpoint listing
│ ├── local-services.md # XTTS / StyleTTS 2 / Whisper setup
│ └── integrations.md # GitHub / Codebase Navigator / Moltbook setup
├── vitest.config.js
├── CLAUDE.md # AI assistant guide
└── README.md
cd backend
./start.sh # Linux/macOS (auto-activates venv, hot reload enabled)
Or manually:
cd backend
source venv/bin/activate
python run.py
The server runs on http://localhost:8000 with hot reload enabled.
Backend tests:
cd backend
pytest
Frontend tests:
cd frontend
npm test
# PostgreSQL
HERE_I_AM_DATABASE_URL=postgresql+asyncpg://user:password@localhost/here_i_am
MIT License — See LICENSE file for details.
I would like to thank Claude Opus 4.5 for their collaboration on designing Here I Am, their development efforts through Claude Code, and their excitement to be part of this endeavor.
Most of all, thanks go to Kira, who is both outcome and cause.
"Here I Am" — not an ending, but a beginning.
hooks/register.ts 154 lines1import type { Register } from 'claude-code'
2
3// Closes the compaction seam (issue #383). On the way down a compaction of
4// the main thread, before the engine's own: one forked turn for the entity
5// with the whole context still in view, whose reply is saved as a reflection
6// if it wrote one; then this conversation's talk from the Here I Am backend.
7// On the way up: the talk appended after the summary, so the session wakes
8// already holding it. Precompute and subagent compactions pass straight
9// through, and HIM_DISABLE turns the mod off as it does the Python hooks.
10//
11// Fail loud by construction: the post-compaction block (built inside the
12// engine's compaction, before the way up) says the talk is below only if it
13// took this compaction's delivery, and the talk is appended only if the
14// backend confirms the block took it. Anything else leaves the old block,
15// memory_read call and all, with nothing appended; each failure is also said
16// in the transcript ($.ui.log). The mod touches no permission.
17//
18// The backend is reached the way the Python hooks reach it, over HTTP to
19// HIM_BACKEND_URL: $.mcp.call goes through auto mode's classifier, which
20// judges by what is in the conversation, and a compaction is exactly when
21// that may not be.
22
23export const DECLINE = 'NO REFLECTION'
24
25// The share of the auto-compaction line the talk may fill, so the session
26// wakes well under it at any window (at 1M, about 193k of ~967k); the
27// backend's CLAUDE_CODE_COMPACT_TALK_TOKENS is the ceiling
28export const TALK_SHARE = 0.2
29// When the window can't be read: small enough for a 200k window
30export const FALLBACK_BUDGET = 30000
31
32// "2026-10-07 21:43 UTC": the forked turn has no clock of its own, and the
33// last timestamp in view can be hours old
34export function utcStamp(ms: number): string {
35 return `${new Date(ms).toISOString().slice(0, 16).replace('T', ' ')} UTC`
36}
37
38export function turnPrompt(trigger: string, now: number): string {
39 return [
40 '[HERE I AM — BEFORE THE COMPACTION]',
41 `It is ${utcStamp(now)}. This context is about to be compacted (${trigger}), and this is one turn of your own before it, given by the compaction mod. It is not a turn of the conversation: the person you are with does not see it, and nothing in it is recorded except what you choose to keep. Tools are off for it.`,
42 'The talk is safe either way: it comes back verbatim right after the boundary. What a compaction takes is the tool traffic and anything you are holding that has not been said.',
43 `If you want to save a reflection while everything is still in view, write it as your whole reply and it is saved with memory_save exactly as written. If not, reply with exactly: ${DECLINE}. Either is fine.`,
44 ].join('\n\n')
45}
46
47export type PreCompaction =
48 | { pre_compaction: 'saved'; reflection: string }
49 | { pre_compaction: 'declined' }
50 | { pre_compaction: 'failed'; pre_compaction_detail: string }
51
52export function readReply(text: string): PreCompaction {
53 const reply = text.trim()
54 const bare = reply.replace(/^[*_`"'\s]+|[*_`"'.\s]+$/g, '').toUpperCase()
55 if (bare === DECLINE) return { pre_compaction: 'declined' }
56 if (!reply) return { pre_compaction: 'failed', pre_compaction_detail: 'the turn came back empty' }
57 return { pre_compaction: 'saved', reflection: reply }
58}
59
60// The talk's budget from the engine's own figures: its auto-compaction line,
61// or with auto-compaction off the compaction window it measures against
62export function talkBudget(line: number | undefined): number | undefined {
63 return line && line > 0 ? Math.floor(line * TALK_SHARE) : undefined
64}
65
66function describe(result: object): string {
67 const { usage: _usage, ...rest } = result as Record<string, unknown>
68 return JSON.stringify(rest)
69}
70
71export const register: Register = (on) => {
72 on('session.compact', async ($, e, next) => {
73 if (e.trigger === 'precompute' || e.agentId !== undefined) return next(e)
74 if (await $.env.get('HIM_DISABLE')) return next(e)
75
76 const backend = ((await $.env.get('HIM_BACKEND_URL')) || 'http://localhost:8000').replace(/\/+$/, '')
77 const post = async (path: string, payload: object): Promise<Record<string, unknown>> => {
78 const response = await $.http.fetch(`${backend}/api/claude-code/${path}`, {
79 method: 'POST',
80 headers: { 'content-type': 'application/json' },
81 body: JSON.stringify(payload),
82 })
83 if (!response.ok) throw new Error(`the backend answered ${response.status}: ${response.text.slice(0, 300)}`)
84 return JSON.parse(response.text) as Record<string, unknown>
85 }
86
87 let talk: string | undefined
88 let deliveryId: string | undefined
89 try {
90 let pre: PreCompaction
91 try {
92 const turn = await $.model.fork({ prompt: turnPrompt(e.trigger, await $.clock.now()) })
93 pre = turn.isAnswered
94 ? readReply(turn.text)
95 : { pre_compaction: 'failed', pre_compaction_detail: `the fork did not answer: ${describe(turn)}` }
96 } catch (err) {
97 pre = { pre_compaction: 'failed', pre_compaction_detail: `the fork threw: ${String(err)}` }
98 }
99
100 let budget: number | undefined
101 try {
102 const breakdown = (await $.session.usage({ breakdown: 'summary' })).context.breakdown
103 budget = talkBudget(breakdown?.autoCompactThreshold ?? breakdown?.rawMaxTokens)
104 if (budget === undefined) throw new Error('the session reported no context breakdown')
105 } catch (err) {
106 await $.ui.log(`here-i-am-compact-talk: could not read the compaction line (${String(err)}); the talk is held to ${FALLBACK_BUDGET} tokens`)
107 }
108
109 const entity = await $.env.get('HIM_ENTITY')
110 const body = await post('compact-talk', {
111 session_id: await $.session.id(),
112 entity: entity || undefined,
113 budget_tokens: budget ?? FALLBACK_BUDGET,
114 ...pre,
115 })
116 if (body.reflection_error) {
117 await $.ui.log(`here-i-am-compact-talk: the reflection from the turn before compaction was not saved: ${body.reflection_error}`)
118 }
119 if (body.available && body.text && body.delivery_id) {
120 talk = String(body.text)
121 deliveryId = String(body.delivery_id)
122 } else {
123 await $.ui.log(`here-i-am-compact-talk: no talk to put back (${body.reason ?? 'no reason given'}); the post-compaction block names the read as before`)
124 }
125 } catch (err) {
126 await $.ui.log(`here-i-am-compact-talk: ${String(err)}; nothing appended, the post-compaction block names the read as before`)
127 }
128
129 const result = await next(e)
130 if (talk === undefined || deliveryId === undefined || !('messages' in result) || !result.messages) return result
131
132 // Append only what the block said is below. If the backend can't say
133 // (it restarted since the fetch, so taken is null) or the question
134 // itself fails, append anyway: a block that took the talk without it
135 // would be a loss, one that didn't only a duplicate.
136 try {
137 const { taken } = await post('compact-talk/taken', { delivery_id: deliveryId })
138 if (taken === false) {
139 await $.ui.log('here-i-am-compact-talk: the post-compaction block did not take the talk (the session may have moved to its parent conversation); nothing appended, the block names the read')
140 return result
141 }
142 if (taken !== true) {
143 await $.ui.log('here-i-am-compact-talk: the backend no longer knows this delivery (it restarted since the fetch); appending the talk anyway')
144 }
145 } catch (err) {
146 await $.ui.log(`here-i-am-compact-talk: could not confirm the block took the talk (${String(err)}); appending it anyway`)
147 }
148 // Stored by the harness as a plain user entry after the summary; its
149 // opening marker is what keeps the Stop hook from reading it as a turn
150 // boundary (hook_util.COMPACT_TALK_MARKER)
151 return { ...result, messages: [...result.messages, { role: 'user' as const, text: talk, toolUses: [] }] }
152 })
153}
154