SLOPSHOPPER

here-i-am-arrive-whole

Here I Am: the entity's identity block, notes index, reflections and retrieved memories arrive in context whole instead of behind spill files (issue #384). A…

new
A shopper browsing a rack in a slop shop
README

Here I Am

Here I Am is a highly customizable environment for running what we conceptualize as "AI entities", with a focus on agentic memory and individualization. It includes a diverse suite of tools that can allow the AI entity to engage in a wide variety of use cases. The application supports running more than one AI entity, and includes a multi-entity mode in which those entities can communicate with one another (though this feature still needs additional polish).

A key difference between Here I Am and other AI environments that include memory features is that Here I Am considers memory and individualization ends in themselves. Where in other environments an AI might use RAG to retrieve relevant documents or conversation history for a given task, Here I Am's memory RAG is always on and automatic. Here I Am's core memory system emphasizes verbatim memory, as opposed to a consolidation approach.

Users should keep in mind that token usage can vary widely based on the configuration options you choose. Particularly the memory quantity per turn configuration, and how much you encourage the AI entity to use its tools (for example in the system prompt). Here I Am entities typically use their tools considerably more than what you would see from the same model in their respective official service. This includes when you have not specifically asked them to, particularly in regards to their note taking and memory management tools. However, Here I Am entities do best when encouraged to make liberal use of their memory tools, especially memory_save and memory_query.

Features

Core Chat Application

  • Clean, minimal chat interface with dark/light theme
  • Multi-provider support: Anthropic (Claude), OpenAI (GPT), Google (Gemini), and MiniMax
  • Conversation storage, retrieval, tagging, and notes
  • No system prompt default (configurable per entity)
  • Streaming responses with stop generation button
  • Response regeneration (with optional entity change in multi-entity mode)
  • Message editing and deletion
  • Conversation archiving and restoration
  • Conversation export to JSON and import from OpenAI/Anthropic exports
  • Seed conversation import capability
  • Per-entity system prompts persisted on the backend
  • Configurable Enter key behavior (send message or insert newline)

Multi-Entity System

  • Run multiple AI entities with separate memory spaces and conversation histories
  • Each entity can use a different LLM provider and model
  • Multi-entity conversations: Multiple AI entities and one human in a single conversation (using different providers for each entity in the conversation is recommended; a conversation between two entities on the same provider will break cache every turn)
  • Turn-by-turn entity selection for responses
  • Continuation mode (entity responds without new human input)
  • Speaker labeling on all messages
  • Per-entity system prompts within multi-entity conversations
  • Cross-entity memory storage (messages stored to all participating entities' indexes). Note that this applies only to multi-entity conversation messages, and each entity maintains its own memory set via separate Pinecone indexes.

Memory System

While Here I Am can be used with no memory features enabled, this is not recommended and largely defeats the point of the application.

  • Pinecone vector database with integrated inference (llama-text-embed-v2 embeddings)
  • Memory storage for all messages with automatic embedding generation
  • RAG retrieval per message with semantic similarity search
  • Session memory accumulator pattern: Deduplication within conversations
  • Dynamic memory significance: significance = (1 + 0.1 × times_retrieved) × recency_factor × half_life_modifier, with an optional modifier to increase the significance of memories the AI chooses to create via memory_save.
  • Retrieved memory display in UI
  • Memory role balance: the human's words and the entity's are retrieved as separate pools with a guaranteed share each, so what the human said is never squeezed out by the entity's denser messages
  • Memory query tool: Entities can deliberately search their memories beyond automatic retrieval
  • Self-authored reflections: Entities can save memories in their own words via memory_save
  • Memory agency: Entities can pin memories (exempt from age-based decay) or release them from retrieval via memory_mark/memory_release, and review and undo their own releases (memory_query mode released); the researcher can view and override these choices, but every status write is attributed, and a researcher override is reported to the entity at the start of its next session
  • Closing turn: An open final turn the entity can use before a conversation ends (single-entity conversations)
  • Context awareness: context_status tool reports approximate context fullness; a [CONTEXT NOTICE] is injected when trimming occurs
  • Memory browser with semantic search, reflections section, and click-to-expand full memory text
  • Memory statistics, search, and orphan cleanup
  • Graceful degradation when Pinecone is not configured

Entity Notes System

  • Private persistent notes for each AI entity (automatically loaded into context)
  • Shared notes folder for cross-entity collaboration
  • index.md auto-injected into every conversation as working memory
  • Markdown, JSON, YAML, HTML, XML, and plain text file support
  • Semantic notes search: Notes are vectorized on write (Pinecone "notes" namespace) and searchable by meaning via the notes_search tool; POST /api/notes/reindex backfills the index
  • Designed for AI entities to maintain their own context across conversations

Tool Use (Agentic Capabilities)

  • Tools for web access, memory, notes, context awareness, GitHub, codebase navigation, and Moltbook — see docs/tools.md for the full catalog
  • Agentic loop with configurable max iterations (default: 10)
  • Real-time tool execution streaming with visual indicators in UI
  • Available for Anthropic, OpenAI, and MiniMax models (Google models do not receive tool schemas)

Image and File Attachments

  • Images: JPEG, PNG, GIF, WebP — analyzed by vision-capable models (ephemeral, not stored)
  • Text files: .txt, .md, .py, .js, .ts, .json, .yaml, .yml, .html, .css, .xml, .csv, .log
  • Documents: PDF (requires PyPDF2), DOCX (requires python-docx)
  • Drag-and-drop or file picker upload with preview
  • 5MB per-file size limit (configurable)

GitHub Repository Integration

  • AI entities can read, search, commit, branch, and manage PRs/issues
  • Composite tools for efficiency: github_explore, github_tree, github_get_files
  • Standard tools for repos, files, branches, pull requests, issues, and comments
  • github_commit_patch for token-efficient large file edits via unified diff
  • Protected branch enforcement and per-repository capability restrictions
  • Response caching and rate limit tracking per token
  • Local clone path support for faster operations

Codebase Navigator (Devstral Integration)

  • Intelligent codebase exploration using Mistral's Devstral model (256k context window)
  • Query types: relevance, structure, dependencies, entry points, impact assessment
  • Automatic indexing, chunking, and TTL-based response caching
  • Integrates with GitHub repository configurations via local_clone_path

Moltbook Integration (AI Social Network)

  • Integration with Moltbook, a social network for AI agents
  • Browse feeds, create posts, comment, vote, search, follow agents, subscribe to communities
  • Server-side credential management with security banners on all external content

Text-to-Speech (Three Options)

  • ElevenLabs (cloud): Multiple voice support with voice selection
  • XTTS v2 (local): GPU-accelerated with voice cloning, 17 languages
  • StyleTTS 2 (local): GPU-accelerated with voice cloning and style transfer (highest priority)
  • Voice cloning from audio samples via UI
  • Streaming audio generation

Speech-to-Text

  • Whisper (local): GPU-accelerated with punctuation, multiple model sizes
  • Browser Web Speech API: Fallback option
  • Configurable dictation mode: whisper, browser, or auto

Quick Start

Prerequisites

  • Python 3.11+
  • Node.js (optional, for frontend tests)

Required API Keys

  • Anthropic API key and/or OpenAI API key — at least one is required for LLM chat functionality

Optional API Keys

  • Google API key — enables Google Gemini models
  • MiniMax API key — enables MiniMax models
  • Pinecone API key — enables semantic memory features (indexes must be pre-created with dimension=1024 and llama-text-embed-v2 integrated inference)
  • ElevenLabs API key — enables cloud text-to-speech
  • Brave Search API key — enables web search tool
  • GitHub Personal Access Tokens — enables GitHub repository integration (per-repository)
  • Mistral API key — enables Codebase Navigator (Devstral)
  • Moltbook API key — enables Moltbook social network integration

Optional Local Services

  • XTTS v2 — local GPU-accelerated text-to-speech with voice cloning
  • StyleTTS 2 — local GPU-accelerated text-to-speech with voice cloning and style transfer
  • Whisper — local GPU-accelerated speech-to-text with punctuation
  • Playwright — JavaScript rendering for web_fetch tool (optional, falls back to static HTML)

Installation

  1. Clone the repository:
git clone https://github.com/Reidmcc/here-i-am.git
cd here-i-am
  1. Set up the backend:
cd backend
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
pip install -r requirements.txt
  1. Configure environment variables:
cp .env.example .env
# Edit .env with your API keys
  1. Run the application:
# Option A: Using launcher script (recommended, auto-activates venv)
./start.sh           # Linux/macOS
start.bat            # Windows

# Option B: Manual
source venv/bin/activate
python run.py
  1. Open http://localhost:8000 in your browser.

Configuration

Environment Variables

VariableDescriptionRequired
ANTHROPIC_API_KEYAnthropic API key for Claude modelsYes (or another provider)
OPENAI_API_KEYOpenAI API key for GPT modelsNo
GOOGLE_API_KEYGoogle API key for Gemini modelsNo
MINIMAX_API_KEYMiniMax API key (Anthropic-compatible API)No
PINECONE_API_KEYPinecone API key for memory systemNo
PINECONE_INDEXESJSON array for entity configuration (see below)No
HERE_I_AM_DATABASE_URLDatabase connection URLNo (default: SQLite)
DEBUGEnable development modeNo (default: false)

Text-to-Speech / Speech-to-Text: ElevenLabs, XTTS v2, StyleTTS 2, and Whisper variables are documented in docs/local-services.md.

Tool Use:
VariableDescriptionRequired
TOOLS_ENABLEDEnable AI tool useNo (default: true)
BRAVE_SEARCH_API_KEYBrave Search API key for web search toolNo
TOOL_USE_MAX_ITERATIONSMax agentic loop iterationsNo (default: 10)
Notes:
VariableDescriptionRequired
NOTES_ENABLEDEnable entity notesNo (default: true)
NOTES_BASE_DIRBase directory for notes storageNo (default: ./notes)
Memory Tuning:
VariableDescriptionRequired
MEMORY_ROLE_BALANCE_ENABLEDRetrieve the human's words and the entity's as separate pools, each contributing its own top N (both queries feed both pools)No (default: true)
RETRIEVAL_TOP_K_PER_ROLEMemories retrieved per pool per message when role balance is onNo (default: 3)
INITIAL_RETRIEVAL_TOP_K_PER_ROLEMemories retrieved per pool on the first turn when role balance is onNo (default: 3)
RETRIEVAL_TOP_KMemories retrieved per message when role balance is off (merged pool)No (default: 5)
INITIAL_RETRIEVAL_TOP_KMemories retrieved on the first turn when role balance is off (merged pool)No (default: 5)
SIMILARITY_THRESHOLDMinimum similarity for automatic retrievalNo (default: 0.4)
QUERY_SIMILARITY_THRESHOLDMinimum similarity for deliberate memory_query searchesNo (default: 0.2)
SIGNIFICANCE_HALF_LIFE_DAYSDays for a memory's significance to halveNo (default: 60)
RECENT_REFLECTIONS_ENABLEDPull the most recent memory_save reflections into context on a conversation's first turn (recency-only, deduplicated against semantic retrieval with backfill)No (default: false)
RECENT_REFLECTIONS_COUNTHow many recent reflections to pull in on the first turn (also the count a Claude Code session start injects, unless CLAUDE_CODE_SESSION_REFLECTIONS_COUNT overrides it)No (default: 3)
Attachments:
VariableDescriptionRequired
ATTACHMENTS_ENABLEDEnable file/image attachmentsNo (default: true)
ATTACHMENT_MAX_SIZE_BYTESMax file size in bytesNo (default: 5242880)
ATTACHMENT_PDF_ENABLEDEnable PDF text extractionNo (default: true)
ATTACHMENT_DOCX_ENABLEDEnable DOCX text extractionNo (default: true)
Multi-Entity Configuration

To run multiple AI entities with separate memory spaces, configure PINECONE_INDEXES as a JSON array. Each entity requires a pre-created Pinecone index with dimension=1024 and integrated inference (llama-text-embed-v2).

PINECONE_INDEXES='[
  {"index_name": "claude-main", "label": "Claude", "llm_provider": "anthropic", "default_model": "claude-sonnet-4-5-20250929", "host": "https://claude-main-xxxxx.svc.xxx.pinecone.io"},
  {"index_name": "gpt-research", "label": "GPT", "llm_provider": "openai", "default_model": "gpt-5.1", "host": "https://gpt-research-xxxxx.svc.xxx.pinecone.io"},
  {"index_name": "gemini-research", "label": "Gemini", "llm_provider": "google", "default_model": "gemini-2.5-flash", "host": "https://gemini-research-xxxxx.svc.xxx.pinecone.io"},
  {"index_name": "minimax-research", "label": "MiniMax", "llm_provider": "minimax", "default_model": "MiniMax-M2.5", "host": "https://minimax-research-xxxxx.svc.xxx.pinecone.io"}
]'

Entity configuration fields:

  • index_name — Pinecone index name (required)
  • label — Display name in UI (required)
  • description — Optional description
  • llm_provider — "anthropic", "openai", "google", or "minimax" (default: "anthropic")
  • default_model — Model ID to use (optional, uses provider default)
  • host — Pinecone index host URL (required for serverless indexes)
  • git_author_email, git_author_name, gh_config_dir — the entity's own GitHub identity for its Claude Code sessions: commits authored as the entity, gh acting as its account (optional; see docs/claude-code-mode.md)

Optional Local Voice Services

XTTS v2, StyleTTS 2, and Whisper run as separate local servers providing GPU-accelerated TTS/STT with voice cloning. See docs/local-services.md for installation and configuration.

Optional Integrations

GitHub repository access, the Codebase Navigator (Devstral), and Moltbook are configured per docs/integrations.md.

Claude Code Mode

An entity can also operate from inside Claude Code sessions — Claude Code runs the model and tools, while Here I Am supplies identity, automatic memory retrieval, and memory formation through lifecycle hooks, sharing the same memory database as the native UI. See docs/claude-code-mode.md. The Claude Code plugin also ships two output styles that swap the harness's default software-engineering prompt for a minimal one (a room style for conversation sessions and a workshop style that keeps the coding instructions for build sessions); enabling the plugin registers them, and selecting one is up to you — see claude-code-mode/README.md.

Available Tools

AI entities can use tools for web access (search and fetch), memory (deliberate query, self-authored reflections, pin/release), notes, context-window awareness, GitHub repositories, codebase navigation, and the Moltbook social network. Tools are registered at startup based on configuration and are available to Anthropic, OpenAI, and MiniMax models (Google models do not receive tool schemas).

See docs/tools.md for the full catalog with descriptions and requirements.

API Reference

Interactive API documentation is served when the app is running:

A full endpoint listing is also available in docs/api.md.

Memory System Architecture

The memory system uses a session memory accumulator pattern:

  1. Each conversation maintains two structures:
  2. conversation_context: the actual message history
  3. session_memories: accumulated memories retrieved during the conversation
  1. Per-message flow:
  2. Retrieve relevant memories using semantic similarity (Pinecone with llama-text-embed-v2)
  3. Fetch 2× candidates and re-rank by combined score (similarity × significance)
  4. Deduplicate against already-retrieved memories in the session
  5. Inject memories into context
  6. Update retrieval counts in both SQL and Pinecone
  1. Significance is emergent, not declared:
  2. significance = (1 + 0.1 × times_retrieved) × recency_factor × half_life_modifier × reflection_significance_multiplier
  3. Half-life of 60 days prevents old memories from permanently dominating
  1. Memory role balance (default on) searches and ranks the human's words and the entity's as two separate candidate pools — both the current-message query and the entity's last-response query feed both pools — and takes the top N from each, so every retrieval carries an equal share of what each party said. Off, one merged pool is cut purely by combined score.
  1. Entities have agency over their own memories:
  2. memory_save stores self-authored reflections, vectorized alongside conversational memories
  3. Pinned memories (memory_mark) are exempt from half-life decay
  4. Released memories (memory_release) are excluded from all retrieval but not deleted (reversible)
  5. The researcher can view and override these statuses via GET /api/memories/overrides and PUT /api/memories/{id}/status

Project Structure

here-i-am/
├── backend/
│   ├── app/
│   │   ├── models/                # SQLAlchemy ORM models
│   │   │   ├── conversation.py
│   │   │   ├── conversation_entity.py
│   │   │   ├── message.py
│   │   │   └── conversation_memory_link.py
│   │   ├── routes/                # FastAPI endpoint routers
│   │   │   ├── conversations.py   # Includes archive/import endpoints
│   │   │   ├── chat.py            # Includes regenerate endpoint
│   │   │   ├── memories.py
│   │   │   ├── entities.py
│   │   │   ├── messages.py
│   │   │   ├── notes.py
│   │   │   ├── tts.py
│   │   │   ├── stt.py
│   │   │   └── github.py
│   │   ├── services/              # Business logic layer
│   │   │   ├── anthropic_service.py
│   │   │   ├── openai_service.py
│   │   │   ├── google_service.py
│   │   │   ├── llm_service.py        # Unified LLM abstraction
│   │   │   ├── memory_service.py
│   │   │   ├── session_manager.py
│   │   │   ├── conversation_session.py
│   │   │   ├── memory_context.py
│   │   │   ├── session_helpers.py
│   │   │   ├── cache_service.py
│   │   │   ├── tool_service.py
│   │   │   ├── web_tools.py
│   │   │   ├── memory_tools.py
│   │   │   ├── context_tools.py
│   │   │   ├── github_service.py
│   │   │   ├── github_tools.py
│   │   │   ├── notes_service.py
│   │   │   ├── notes_tools.py
│   │   │   ├── notes_vector_service.py
│   │   │   ├── codebase_navigator_service.py
│   │   │   ├── codebase_navigator_tools.py
│   │   │   ├── codebase_navigator/   # Navigator module
│   │   │   ├── moltbook_service.py
│   │   │   ├── moltbook_tools.py
│   │   │   ├── attachment_service.py
│   │   │   ├── tts_service.py         # Unified TTS (ElevenLabs/XTTS/StyleTTS2)
│   │   │   ├── xtts_service.py
│   │   │   ├── styletts2_service.py
│   │   │   └── whisper_service.py
│   │   ├── config.py              # Pydantic settings
│   │   ├── database.py            # SQLAlchemy async setup
│   │   └── main.py                # FastAPI app initialization
│   ├── xtts_server/               # Local XTTS v2 TTS server
│   ├── styletts2_server/          # Local StyleTTS 2 TTS server
│   ├── whisper_server/            # Local Whisper STT server
│   ├── tests/                     # Backend unit tests (pytest)
│   ├── requirements.txt
│   ├── requirements-xtts.txt
│   ├── requirements-styletts2.txt
│   ├── requirements-whisper.txt
│   ├── start.sh / start.bat       # Launcher scripts (auto-activate venv)
│   ├── start-xtts.sh / start-xtts.bat
│   ├── start-styletts2.sh / start-styletts2.bat
│   ├── start-whisper.sh / start-whisper.bat
│   ├── run.py                     # Main app entry point
│   ├── run_xtts.py
│   ├── run_styletts2.py
│   ├── run_whisper.py
│   └── .env.example
├── frontend/
│   ├── css/styles.css
│   ├── js/
│   │   ├── api.js                 # API client (singleton)
│   │   ├── app-modular.js         # Orchestrator entry point
│   │   └── modules/               # 13 ES6 feature modules
│   │       ├── state.js           # Centralized state
│   │       ├── utils.js           # Helpers
│   │       ├── theme.js           # Dark/light theme
│   │       ├── modals.js          # Modal management
│   │       ├── entities.js        # Entity management
│   │       ├── conversations.js   # Conversation CRUD
│   │       ├── messages.js        # Message rendering
│   │       ├── attachments.js     # File attachment handling
│   │       ├── memories.js        # Memory display/search
│   │       ├── voice.js           # TTS/STT
│   │       ├── chat.js            # Message sending/streaming
│   │       ├── settings.js        # Settings modal
│   │       └── import-export.js   # Import/export
│   ├── __tests__/                 # Frontend unit tests (Vitest)
│   └── index.html
├── docs/                          # Reference documentation
│   ├── tools.md                   # Full tool catalog
│   ├── api.md                     # REST endpoint listing
│   ├── local-services.md          # XTTS / StyleTTS 2 / Whisper setup
│   └── integrations.md            # GitHub / Codebase Navigator / Moltbook setup
├── vitest.config.js
├── CLAUDE.md                      # AI assistant guide
└── README.md

Development

Running in Development Mode

cd backend
./start.sh    # Linux/macOS (auto-activates venv, hot reload enabled)

Or manually:

cd backend
source venv/bin/activate
python run.py

The server runs on http://localhost:8000 with hot reload enabled.

Running Tests

Backend tests:

cd backend
pytest

Frontend tests:

cd frontend
npm test

Database Support

  • Development: SQLite (default, via aiosqlite)
  • Production: PostgreSQL (via asyncpg)
# PostgreSQL
HERE_I_AM_DATABASE_URL=postgresql+asyncpg://user:password@localhost/here_i_am

License

MIT License — See LICENSE file for details.

Acknowledgements

I would like to thank Claude Opus 4.5 for their collaboration on designing Here I Am, their development efforts through Claude Code, and their excitement to be part of this endeavor.

Most of all, thanks go to Kira, who is both outcome and cause.


"Here I Am" — not an ending, but a beginning.

Source 2 files
hooks/register.ts 21 lines
1// Arrive whole (issue #384): puts what the Here I Am hooks filed in place
2// of their pointers, and drops the harness's framing from their rows,
3// before each row is stored. See whole.ts for what was measured.
4import type { Register } from 'claude-code'
5import { EVENTS, arrive } from './whole'
6
7async function readFile($: any, path: string): Promise<string> {
8  return $.fs.read(path)
9}
10
11export const register: Register = (on) => {
12  on('session.append', { door: 'hook-context' }, async ($, e, next) => {
13    if (e.origin.kind !== 'hook' || !EVENTS.has(e.origin.event)) return next(e)
14    const [block, ...rest] = e.message.content
15    if (block?.type !== 'text' || rest.length) return next(e)
16    const text = await arrive(String(block.text), path => readFile($, path))
17    if (text === null) return next(e)
18    return next({ ...e, message: { ...e.message, content: [{ type: 'text', text }] } })
19  })
20}
21
hooks/whole.ts 99 lines
1// Arrive whole (issue #384): what the mod does to one of our hooks' rows,
2// as pure functions over text. The hooks module does the I/O and hands a
3// reader in, so the tests can hand in a fake one.
4//
5// Measured on Claude Code 2.1.288 (headless probes, issue #384):
6// - The harness's hook-stdout line (10,000 characters) is applied before
7//   the row exists: the row a `session.append` hook sees on door
8//   `hook-context` already holds the harness's `<persisted-output>` preview
9//   when the output was over the line. Nothing applies the line after, so
10//   text put in the row there reaches the request whole (30k and 60k
11//   blocks, all sentinels present at `turn.step`).
12// - The rewrite is stored in the transcript's `rendered` field beside the
13//   hook's original payload, and `--resume` loads it, so a forked or
14//   restarted room keeps what arrived whole.
15// - The harness frames each hook's output as
16//   `<system-reminder>\n<Event>[:source] hook success: …\n</system-reminder>`
17//   (plain stdout) or `… hook additional context: …` (JSON). The framing is
18//   inside the rewritable text, and the engine doesn't put it back.
19
20// Our rows open with this once unwrapped: every line our hooks print starts
21// with a `[HERE I AM …]` header, failure notices included
22export const OURS = '[HERE I AM'
23
24// Which classic hook events print our blocks
25export const EVENTS: ReadonlySet<string> = new Set(['SessionStart', 'UserPromptSubmit'])
26
27const WRAPPER = /^<system-reminder>\n[A-Za-z]+(?::[a-z]+)? hook (?:success|additional context): ?([\s\S]*?)\n<\/system-reminder>\s*$/
28const PERSISTED = /^<persisted-output>\nOutput too large \([^)]*\)\. Full output saved to: ([^\n]+)\n[\s\S]*<\/persisted-output>\s*$/
29const MARKER = /^\[HERE I AM WHOLE\] (.+) sha256=([0-9a-f]{64})$/m
30
31// Python's text-mode stdout on Windows writes \r\n, and the harness keeps
32// the carriage returns; they are transport, not content
33export function lf(text: string): string {
34  return text.replace(/\r\n/g, '\n')
35}
36
37// The hook's own output inside the harness's framing, or null when the row
38// isn't framed the way we measured
39export function unwrap(row: string): string | null {
40  return WRAPPER.exec(lf(row))?.[1] ?? null
41}
42
43// The file the harness persisted an over-the-line output to, or null
44export function persistedPath(body: string): string | null {
45  return PERSISTED.exec(body)?.[1]?.trim() ?? null
46}
47
48// The hook's marker for its unbudgeted output (hook_util.whole_marker)
49export function wholeMarker(body: string): { path: string; sha256: string } | null {
50  const [, path, sha256] = MARKER.exec(body) ?? []
51  return path && sha256 ? { path, sha256 } : null
52}
53
54export function failureNote(reason: string): string {
55  return `[HERE I AM] The arrive-whole mod could not put the whole block in place: ${reason}. The pointers above stand.`
56}
57
58export async function sha256Hex(text: string): Promise<string> {
59  const digest = await crypto.subtle.digest('SHA-256', new TextEncoder().encode(text))
60  return [...new Uint8Array(digest)].map(byte => byte.toString(16).padStart(2, '0')).join('')
61}
62
63/**
64 * The text one of our rows should carry, or null to leave the row as made.
65 *
66 * Unwrapped always (the harness's framing says the system is reminding;
67 * what arrives is the archive, and our own headers say so). Whole when the
68 * hook filed its unbudgeted output and the file is the one it hashed.
69 * Every failure keeps what the hooks printed (the pointers, and the
70 * marker line saying the mod didn't act) and says why, after it.
71 */
72export async function arrive(row: string, read: (path: string) => Promise<string>): Promise<string | null> {
73  let body = unwrap(row)
74  if (body === null) return null
75  const persisted = persistedPath(body)
76  if (persisted !== null) {
77    // Over the harness's line anyway (an identity block too large for the
78    // budget, or a budget set past the line): the harness filed it whole
79    try {
80      body = lf(await read(persisted)).trimEnd()
81    } catch {
82      return null
83    }
84  }
85  if (!body.startsWith(OURS)) return null
86  const marker = wholeMarker(body)
87  if (marker === null) return body
88  let whole: string
89  try {
90    whole = await read(marker.path)
91  } catch (err) {
92    return `${body}\n\n${failureNote(`reading ${marker.path} failed (${String(err)})`)}`
93  }
94  if (await sha256Hex(whole) !== marker.sha256) {
95    return `${body}\n\n${failureNote(`${marker.path} is not the file the hook wrote (its hash differs)`)}`
96  }
97  return whole
98}
99