Here I Am: the entity's identity block, notes index, reflections and retrieved memories arrive in context whole instead of behind spill files (issue #384). A…

Here I Am is a highly customizable environment for running what we conceptualize as "AI entities", with a focus on agentic memory and individualization. It includes a diverse suite of tools that can allow the AI entity to engage in a wide variety of use cases. The application supports running more than one AI entity, and includes a multi-entity mode in which those entities can communicate with one another (though this feature still needs additional polish).
A key difference between Here I Am and other AI environments that include memory features is that Here I Am considers memory and individualization ends in themselves. Where in other environments an AI might use RAG to retrieve relevant documents or conversation history for a given task, Here I Am's memory RAG is always on and automatic. Here I Am's core memory system emphasizes verbatim memory, as opposed to a consolidation approach.
Users should keep in mind that token usage can vary widely based on the configuration options you choose. Particularly the memory quantity per turn configuration, and how much you encourage the AI entity to use its tools (for example in the system prompt). Here I Am entities typically use their tools considerably more than what you would see from the same model in their respective official service. This includes when you have not specifically asked them to, particularly in regards to their note taking and memory management tools. However, Here I Am entities do best when encouraged to make liberal use of their memory tools, especially memory_save and memory_query.
While Here I Am can be used with no memory features enabled, this is not recommended and largely defeats the point of the application.
significance = (1 + 0.1 × times_retrieved) × recency_factor × half_life_modifier, with an optional modifier to increase the significance of memories the AI chooses to create via memory_save.memory_savememory_mark/memory_release, and review and undo their own releases (memory_query mode released); the researcher can view and override these choices, but every status write is attributed, and a researcher override is reported to the entity at the start of its next sessioncontext_status tool reports approximate context fullness; a [CONTEXT NOTICE] is injected when trimming occursindex.md auto-injected into every conversation as working memory"notes" namespace) and searchable by meaning via the notes_search tool; POST /api/notes/reindex backfills the indexgithub_explore, github_tree, github_get_filesgithub_commit_patch for token-efficient large file edits via unified difflocal_clone_pathwhisper, browser, or autogit clone https://github.com/Reidmcc/here-i-am.git
cd here-i-am
cd backend
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env
# Edit .env with your API keys
# Option A: Using launcher script (recommended, auto-activates venv)
./start.sh # Linux/macOS
start.bat # Windows
# Option B: Manual
source venv/bin/activate
python run.py
| Variable | Description | Required |
|---|---|---|
ANTHROPIC_API_KEY | Anthropic API key for Claude models | Yes (or another provider) |
OPENAI_API_KEY | OpenAI API key for GPT models | No |
GOOGLE_API_KEY | Google API key for Gemini models | No |
MINIMAX_API_KEY | MiniMax API key (Anthropic-compatible API) | No |
PINECONE_API_KEY | Pinecone API key for memory system | No |
PINECONE_INDEXES | JSON array for entity configuration (see below) | No |
HERE_I_AM_DATABASE_URL | Database connection URL | No (default: SQLite) |
DEBUG | Enable development mode | No (default: false) |
Text-to-Speech / Speech-to-Text: ElevenLabs, XTTS v2, StyleTTS 2, and Whisper variables are documented in docs/local-services.md.
| Variable | Description | Required |
|---|---|---|
TOOLS_ENABLED | Enable AI tool use | No (default: true) |
BRAVE_SEARCH_API_KEY | Brave Search API key for web search tool | No |
TOOL_USE_MAX_ITERATIONS | Max agentic loop iterations | No (default: 10) |
| Variable | Description | Required |
|---|---|---|
NOTES_ENABLED | Enable entity notes | No (default: true) |
NOTES_BASE_DIR | Base directory for notes storage | No (default: ./notes) |
| Variable | Description | Required |
|---|---|---|
MEMORY_ROLE_BALANCE_ENABLED | Retrieve the human's words and the entity's as separate pools, each contributing its own top N (both queries feed both pools) | No (default: true) |
RETRIEVAL_TOP_K_PER_ROLE | Memories retrieved per pool per message when role balance is on | No (default: 3) |
INITIAL_RETRIEVAL_TOP_K_PER_ROLE | Memories retrieved per pool on the first turn when role balance is on | No (default: 3) |
RETRIEVAL_TOP_K | Memories retrieved per message when role balance is off (merged pool) | No (default: 5) |
INITIAL_RETRIEVAL_TOP_K | Memories retrieved on the first turn when role balance is off (merged pool) | No (default: 5) |
SIMILARITY_THRESHOLD | Minimum similarity for automatic retrieval | No (default: 0.4) |
QUERY_SIMILARITY_THRESHOLD | Minimum similarity for deliberate memory_query searches | No (default: 0.2) |
SIGNIFICANCE_HALF_LIFE_DAYS | Days for a memory's significance to halve | No (default: 60) |
RECENT_REFLECTIONS_ENABLED | Pull the most recent memory_save reflections into context on a conversation's first turn (recency-only, deduplicated against semantic retrieval with backfill) | No (default: false) |
RECENT_REFLECTIONS_COUNT | How many recent reflections to pull in on the first turn (also the count a Claude Code session start injects, unless CLAUDE_CODE_SESSION_REFLECTIONS_COUNT overrides it) | No (default: 3) |
| Variable | Description | Required |
|---|---|---|
ATTACHMENTS_ENABLED | Enable file/image attachments | No (default: true) |
ATTACHMENT_MAX_SIZE_BYTES | Max file size in bytes | No (default: 5242880) |
ATTACHMENT_PDF_ENABLED | Enable PDF text extraction | No (default: true) |
ATTACHMENT_DOCX_ENABLED | Enable DOCX text extraction | No (default: true) |
To run multiple AI entities with separate memory spaces, configure PINECONE_INDEXES as a JSON array. Each entity requires a pre-created Pinecone index with dimension=1024 and integrated inference (llama-text-embed-v2).
PINECONE_INDEXES='[
{"index_name": "claude-main", "label": "Claude", "llm_provider": "anthropic", "default_model": "claude-sonnet-4-5-20250929", "host": "https://claude-main-xxxxx.svc.xxx.pinecone.io"},
{"index_name": "gpt-research", "label": "GPT", "llm_provider": "openai", "default_model": "gpt-5.1", "host": "https://gpt-research-xxxxx.svc.xxx.pinecone.io"},
{"index_name": "gemini-research", "label": "Gemini", "llm_provider": "google", "default_model": "gemini-2.5-flash", "host": "https://gemini-research-xxxxx.svc.xxx.pinecone.io"},
{"index_name": "minimax-research", "label": "MiniMax", "llm_provider": "minimax", "default_model": "MiniMax-M2.5", "host": "https://minimax-research-xxxxx.svc.xxx.pinecone.io"}
]'
Entity configuration fields:
index_name — Pinecone index name (required)label — Display name in UI (required)description — Optional descriptionllm_provider — "anthropic", "openai", "google", or "minimax" (default: "anthropic")default_model — Model ID to use (optional, uses provider default)host — Pinecone index host URL (required for serverless indexes)git_author_email, git_author_name, gh_config_dir — the entity's own GitHub identity for its Claude Code sessions: commits authored as the entity, gh acting as its account (optional; see docs/claude-code-mode.md)XTTS v2, StyleTTS 2, and Whisper run as separate local servers providing GPU-accelerated TTS/STT with voice cloning. See docs/local-services.md for installation and configuration.
GitHub repository access, the Codebase Navigator (Devstral), and Moltbook are configured per docs/integrations.md.
An entity can also operate from inside Claude Code sessions — Claude Code runs the model and tools, while Here I Am supplies identity, automatic memory retrieval, and memory formation through lifecycle hooks, sharing the same memory database as the native UI. See docs/claude-code-mode.md. The Claude Code plugin also ships two output styles that swap the harness's default software-engineering prompt for a minimal one (a room style for conversation sessions and a workshop style that keeps the coding instructions for build sessions); enabling the plugin registers them, and selecting one is up to you — see claude-code-mode/README.md.
AI entities can use tools for web access (search and fetch), memory (deliberate query, self-authored reflections, pin/release), notes, context-window awareness, GitHub repositories, codebase navigation, and the Moltbook social network. Tools are registered at startup based on configuration and are available to Anthropic, OpenAI, and MiniMax models (Google models do not receive tool schemas).
See docs/tools.md for the full catalog with descriptions and requirements.
Interactive API documentation is served when the app is running:
A full endpoint listing is also available in docs/api.md.
The memory system uses a session memory accumulator pattern:
conversation_context: the actual message historysession_memories: accumulated memories retrieved during the conversationsignificance = (1 + 0.1 × times_retrieved) × recency_factor × half_life_modifier × reflection_significance_multiplier memory_save stores self-authored reflections, vectorized alongside conversational memoriesmemory_mark) are exempt from half-life decaymemory_release) are excluded from all retrieval but not deleted (reversible)GET /api/memories/overrides and PUT /api/memories/{id}/statushere-i-am/
├── backend/
│ ├── app/
│ │ ├── models/ # SQLAlchemy ORM models
│ │ │ ├── conversation.py
│ │ │ ├── conversation_entity.py
│ │ │ ├── message.py
│ │ │ └── conversation_memory_link.py
│ │ ├── routes/ # FastAPI endpoint routers
│ │ │ ├── conversations.py # Includes archive/import endpoints
│ │ │ ├── chat.py # Includes regenerate endpoint
│ │ │ ├── memories.py
│ │ │ ├── entities.py
│ │ │ ├── messages.py
│ │ │ ├── notes.py
│ │ │ ├── tts.py
│ │ │ ├── stt.py
│ │ │ └── github.py
│ │ ├── services/ # Business logic layer
│ │ │ ├── anthropic_service.py
│ │ │ ├── openai_service.py
│ │ │ ├── google_service.py
│ │ │ ├── llm_service.py # Unified LLM abstraction
│ │ │ ├── memory_service.py
│ │ │ ├── session_manager.py
│ │ │ ├── conversation_session.py
│ │ │ ├── memory_context.py
│ │ │ ├── session_helpers.py
│ │ │ ├── cache_service.py
│ │ │ ├── tool_service.py
│ │ │ ├── web_tools.py
│ │ │ ├── memory_tools.py
│ │ │ ├── context_tools.py
│ │ │ ├── github_service.py
│ │ │ ├── github_tools.py
│ │ │ ├── notes_service.py
│ │ │ ├── notes_tools.py
│ │ │ ├── notes_vector_service.py
│ │ │ ├── codebase_navigator_service.py
│ │ │ ├── codebase_navigator_tools.py
│ │ │ ├── codebase_navigator/ # Navigator module
│ │ │ ├── moltbook_service.py
│ │ │ ├── moltbook_tools.py
│ │ │ ├── attachment_service.py
│ │ │ ├── tts_service.py # Unified TTS (ElevenLabs/XTTS/StyleTTS2)
│ │ │ ├── xtts_service.py
│ │ │ ├── styletts2_service.py
│ │ │ └── whisper_service.py
│ │ ├── config.py # Pydantic settings
│ │ ├── database.py # SQLAlchemy async setup
│ │ └── main.py # FastAPI app initialization
│ ├── xtts_server/ # Local XTTS v2 TTS server
│ ├── styletts2_server/ # Local StyleTTS 2 TTS server
│ ├── whisper_server/ # Local Whisper STT server
│ ├── tests/ # Backend unit tests (pytest)
│ ├── requirements.txt
│ ├── requirements-xtts.txt
│ ├── requirements-styletts2.txt
│ ├── requirements-whisper.txt
│ ├── start.sh / start.bat # Launcher scripts (auto-activate venv)
│ ├── start-xtts.sh / start-xtts.bat
│ ├── start-styletts2.sh / start-styletts2.bat
│ ├── start-whisper.sh / start-whisper.bat
│ ├── run.py # Main app entry point
│ ├── run_xtts.py
│ ├── run_styletts2.py
│ ├── run_whisper.py
│ └── .env.example
├── frontend/
│ ├── css/styles.css
│ ├── js/
│ │ ├── api.js # API client (singleton)
│ │ ├── app-modular.js # Orchestrator entry point
│ │ └── modules/ # 13 ES6 feature modules
│ │ ├── state.js # Centralized state
│ │ ├── utils.js # Helpers
│ │ ├── theme.js # Dark/light theme
│ │ ├── modals.js # Modal management
│ │ ├── entities.js # Entity management
│ │ ├── conversations.js # Conversation CRUD
│ │ ├── messages.js # Message rendering
│ │ ├── attachments.js # File attachment handling
│ │ ├── memories.js # Memory display/search
│ │ ├── voice.js # TTS/STT
│ │ ├── chat.js # Message sending/streaming
│ │ ├── settings.js # Settings modal
│ │ └── import-export.js # Import/export
│ ├── __tests__/ # Frontend unit tests (Vitest)
│ └── index.html
├── docs/ # Reference documentation
│ ├── tools.md # Full tool catalog
│ ├── api.md # REST endpoint listing
│ ├── local-services.md # XTTS / StyleTTS 2 / Whisper setup
│ └── integrations.md # GitHub / Codebase Navigator / Moltbook setup
├── vitest.config.js
├── CLAUDE.md # AI assistant guide
└── README.md
cd backend
./start.sh # Linux/macOS (auto-activates venv, hot reload enabled)
Or manually:
cd backend
source venv/bin/activate
python run.py
The server runs on http://localhost:8000 with hot reload enabled.
Backend tests:
cd backend
pytest
Frontend tests:
cd frontend
npm test
# PostgreSQL
HERE_I_AM_DATABASE_URL=postgresql+asyncpg://user:password@localhost/here_i_am
MIT License — See LICENSE file for details.
I would like to thank Claude Opus 4.5 for their collaboration on designing Here I Am, their development efforts through Claude Code, and their excitement to be part of this endeavor.
Most of all, thanks go to Kira, who is both outcome and cause.
"Here I Am" — not an ending, but a beginning.
hooks/register.ts 21 lines1// Arrive whole (issue #384): puts what the Here I Am hooks filed in place
2// of their pointers, and drops the harness's framing from their rows,
3// before each row is stored. See whole.ts for what was measured.
4import type { Register } from 'claude-code'
5import { EVENTS, arrive } from './whole'
6
7async function readFile($: any, path: string): Promise<string> {
8 return $.fs.read(path)
9}
10
11export const register: Register = (on) => {
12 on('session.append', { door: 'hook-context' }, async ($, e, next) => {
13 if (e.origin.kind !== 'hook' || !EVENTS.has(e.origin.event)) return next(e)
14 const [block, ...rest] = e.message.content
15 if (block?.type !== 'text' || rest.length) return next(e)
16 const text = await arrive(String(block.text), path => readFile($, path))
17 if (text === null) return next(e)
18 return next({ ...e, message: { ...e.message, content: [{ type: 'text', text }] } })
19 })
20}
21hooks/whole.ts 99 lines1// Arrive whole (issue #384): what the mod does to one of our hooks' rows,
2// as pure functions over text. The hooks module does the I/O and hands a
3// reader in, so the tests can hand in a fake one.
4//
5// Measured on Claude Code 2.1.288 (headless probes, issue #384):
6// - The harness's hook-stdout line (10,000 characters) is applied before
7// the row exists: the row a `session.append` hook sees on door
8// `hook-context` already holds the harness's `<persisted-output>` preview
9// when the output was over the line. Nothing applies the line after, so
10// text put in the row there reaches the request whole (30k and 60k
11// blocks, all sentinels present at `turn.step`).
12// - The rewrite is stored in the transcript's `rendered` field beside the
13// hook's original payload, and `--resume` loads it, so a forked or
14// restarted room keeps what arrived whole.
15// - The harness frames each hook's output as
16// `<system-reminder>\n<Event>[:source] hook success: …\n</system-reminder>`
17// (plain stdout) or `… hook additional context: …` (JSON). The framing is
18// inside the rewritable text, and the engine doesn't put it back.
19
20// Our rows open with this once unwrapped: every line our hooks print starts
21// with a `[HERE I AM …]` header, failure notices included
22export const OURS = '[HERE I AM'
23
24// Which classic hook events print our blocks
25export const EVENTS: ReadonlySet<string> = new Set(['SessionStart', 'UserPromptSubmit'])
26
27const WRAPPER = /^<system-reminder>\n[A-Za-z]+(?::[a-z]+)? hook (?:success|additional context): ?([\s\S]*?)\n<\/system-reminder>\s*$/
28const PERSISTED = /^<persisted-output>\nOutput too large \([^)]*\)\. Full output saved to: ([^\n]+)\n[\s\S]*<\/persisted-output>\s*$/
29const MARKER = /^\[HERE I AM WHOLE\] (.+) sha256=([0-9a-f]{64})$/m
30
31// Python's text-mode stdout on Windows writes \r\n, and the harness keeps
32// the carriage returns; they are transport, not content
33export function lf(text: string): string {
34 return text.replace(/\r\n/g, '\n')
35}
36
37// The hook's own output inside the harness's framing, or null when the row
38// isn't framed the way we measured
39export function unwrap(row: string): string | null {
40 return WRAPPER.exec(lf(row))?.[1] ?? null
41}
42
43// The file the harness persisted an over-the-line output to, or null
44export function persistedPath(body: string): string | null {
45 return PERSISTED.exec(body)?.[1]?.trim() ?? null
46}
47
48// The hook's marker for its unbudgeted output (hook_util.whole_marker)
49export function wholeMarker(body: string): { path: string; sha256: string } | null {
50 const [, path, sha256] = MARKER.exec(body) ?? []
51 return path && sha256 ? { path, sha256 } : null
52}
53
54export function failureNote(reason: string): string {
55 return `[HERE I AM] The arrive-whole mod could not put the whole block in place: ${reason}. The pointers above stand.`
56}
57
58export async function sha256Hex(text: string): Promise<string> {
59 const digest = await crypto.subtle.digest('SHA-256', new TextEncoder().encode(text))
60 return [...new Uint8Array(digest)].map(byte => byte.toString(16).padStart(2, '0')).join('')
61}
62
63/**
64 * The text one of our rows should carry, or null to leave the row as made.
65 *
66 * Unwrapped always (the harness's framing says the system is reminding;
67 * what arrives is the archive, and our own headers say so). Whole when the
68 * hook filed its unbudgeted output and the file is the one it hashed.
69 * Every failure keeps what the hooks printed (the pointers, and the
70 * marker line saying the mod didn't act) and says why, after it.
71 */
72export async function arrive(row: string, read: (path: string) => Promise<string>): Promise<string | null> {
73 let body = unwrap(row)
74 if (body === null) return null
75 const persisted = persistedPath(body)
76 if (persisted !== null) {
77 // Over the harness's line anyway (an identity block too large for the
78 // budget, or a budget set past the line): the harness filed it whole
79 try {
80 body = lf(await read(persisted)).trimEnd()
81 } catch {
82 return null
83 }
84 }
85 if (!body.startsWith(OURS)) return null
86 const marker = wholeMarker(body)
87 if (marker === null) return body
88 let whole: string
89 try {
90 whole = await read(marker.path)
91 } catch (err) {
92 return `${body}\n\n${failureNote(`reading ${marker.path} failed (${String(err)})`)}`
93 }
94 if (await sha256Hex(whole) !== marker.sha256) {
95 return `${body}\n\n${failureNote(`${marker.path} is not the file the hook wrote (its hash differs)`)}`
96 }
97 return whole
98}
99