Skip to content

Observe — API reference

CLI

main

main()

Entry point for mc-hook-observe.

Transcript

format_for_extraction

format_for_extraction(
    messages: list[dict], max_chars: int = 60000
) -> str

Format filtered messages into readable text for the extraction prompt.

PARAMETER DESCRIPTION
messages

Filtered transcript message dicts.

TYPE: list[dict]

max_chars

Maximum length of the returned string. If exceeded, the leading portion is dropped so the most recent content is kept.

TYPE: int DEFAULT: 60000

RETURNS DESCRIPTION
str

A single string with one ROLE: text line per message.

read_new_lines

read_new_lines(
    transcript_path: str, byte_offset: int
) -> tuple[list[dict], int, dict]

Read new JSONL transcript lines from byte_offset to EOF.

Only returns user and assistant messages, skipping tool_result and tool_use blocks.

PARAMETER DESCRIPTION
transcript_path

Filesystem path to the JSONL transcript.

TYPE: str

byte_offset

Last-read offset; reading resumes from here.

TYPE: int

RETURNS DESCRIPTION
list[dict]

A tuple (filtered_messages, new_byte_offset, metadata).

int

metadata contains project and branch derived from the

dict

transcript entries' cwd and gitBranch fields; the last

tuple[list[dict], int, dict]

non-empty values encountered are used.

Cursor

get_cursor

get_cursor(transcript_path: str) -> int

Return the byte offset for a transcript path, or 0 if unseen.

load_cursors

load_cursors() -> dict

Load the cursor file, returning {} on missing or corrupt input.

save_cursors

save_cursors(cursors: dict) -> None

Atomically write the cursors dict, pruning stale entries.

set_cursor

set_cursor(transcript_path: str, offset: int) -> None

Advance the byte offset for a transcript path and persist the cursors file.

Writing a fresh record also clears any extraction-failure state recorded by :func:record_failure, since advancing means moving past the window that was failing.

Extraction

build_extraction_prompt

build_extraction_prompt(
    conversation_text: str,
    *,
    project: str | None = None,
    branch: str | None = None,
) -> str

Wrap conversation text in the extraction system prompt.

PARAMETER DESCRIPTION
conversation_text

Formatted conversation transcript.

TYPE: str

project

Project identifier to include in the context header.

TYPE: str | None DEFAULT: None

branch

Branch name to include in the context header.

TYPE: str | None DEFAULT: None

RETURNS DESCRIPTION
str

The full extraction prompt ready to pass to the LLM.

call_haiku

call_haiku(prompt: str) -> str | None

Call the configured extraction model via the shared LLM module.

The model is the extraction-model setting (default haiku); the timeout follows the llm-timeout setting.

PARAMETER DESCRIPTION
prompt

The full extraction prompt to send.

TYPE: str

RETURNS DESCRIPTION
str | None

Stripped stdout from the model, or None on failure.

extract_observations

extract_observations(
    conversation_text: str,
    *,
    project: str | None = None,
    branch: str | None = None,
) -> dict | None

Run the full extraction pipeline: build prompt, call Haiku, parse.

PARAMETER DESCRIPTION
conversation_text

Formatted conversation transcript.

TYPE: str

project

Project identifier passed into the prompt header.

TYPE: str | None DEFAULT: None

branch

Branch name passed into the prompt header.

TYPE: str | None DEFAULT: None

RETURNS DESCRIPTION
dict | None

None if the LLM call itself failed (timeout / non-zero exit) — a

dict | None

transient failure the caller should retry by leaving the cursor put.

dict | None

{} if the model responded but its output could not be parsed (the

dict | None

call succeeded, so treat the window as handled). The parsed extraction

dict | None

dict on success — its arrays may be empty when nothing was notable.

parse_extraction

parse_extraction(raw_output: str) -> dict | None

Parse JSON from Haiku's response, tolerating markdown code fences.

PARAMETER DESCRIPTION
raw_output

The raw stdout returned by Haiku.

TYPE: str

RETURNS DESCRIPTION
dict | None

The parsed extraction dict, or None if the response cannot be

dict | None

parsed into a dict containing any of the expected keys.

Scope

resolve_scope

resolve_scope(
    item: dict,
    category: str,
    project: str | None,
    branch: str | None,
) -> str

Determine the scope for an extracted observation item.

Strategy: trust the LLM's scope classification when it is valid, then fall back to per-category heuristics.

PARAMETER DESCRIPTION
item

Extracted item dict (may contain "scope" from LLM).

TYPE: dict

category

Extraction category (facts, preferences, decisions, procedures, bugs_fixed).

TYPE: str

project

Current project identifier, or None.

TYPE: str | None

branch

Current branch name, or None.

TYPE: str | None

RETURNS DESCRIPTION
str

One of "global", "project", "branch".

Storage

store_entities_and_relationships

store_entities_and_relationships(extraction: dict) -> None

Upsert entities and relationships from an extraction dict.

Entities referenced by relationships but absent from the entities list are auto-created with type "unknown".

PARAMETER DESCRIPTION
extraction

Parsed extraction dict with entities and relationships keys.

TYPE: dict

store_extraction

store_extraction(
    extraction: dict,
    *,
    project: str | None = None,
    branch: str | None = None,
) -> None

Store extracted facts, preferences, decisions, and bugs to memory.db.

PARAMETER DESCRIPTION
extraction

Parsed extraction dict with category keys.

TYPE: dict

project

Project identifier used to resolve scope and assign entries.

TYPE: str | None DEFAULT: None

branch

Branch name used to resolve scope and assign entries.

TYPE: str | None DEFAULT: None

write_raw_extraction

write_raw_extraction(
    extraction: dict,
    trigger: str,
    *,
    project: str | None = None,
    branch: str | None = None,
) -> None

Append an extraction record to the day's JSONL log file.

PARAMETER DESCRIPTION
extraction

Parsed extraction dict returned by the LLM.

TYPE: dict

trigger

The hook trigger name (e.g. "precompact").

TYPE: str

project

Project identifier captured from the transcript.

TYPE: str | None DEFAULT: None

branch

Branch name captured from the transcript.

TYPE: str | None DEFAULT: None