~kris/9p

llm9p

llm9p/internal/llm/session.go -rw-r--r-- 20.8 KiB
b96d9137 — pdfinn 6 months ago
fix(llm): seed session model from backend and persist sessions across connections

Two bugs found during gpt-oss/Ollama operational test:

1. DefaultSessionDefaults() hardcoded the Claude model ID, causing
   llm9p to send "claude-sonnet-4-5-20250929" to Ollama which has no
   such model. NewSessionManager now seeds defaults.Model from
   apiClient.Model() so the session inherits the backend's configured
   model (e.g. "gpt-oss:20b").

2. Sessions started with refs=0 and were auto-deleted the moment the
   first 9P fid using them was clunked. The plan9port 9p CLI tool
   opens a new connection per command, so a session created by
   "9p read new" was gone before the next "9p write N/ask" could use
   it. Sessions now start with refs=1 (the session holds a reference
   to itself) and persist until explicitly closed via "ctl close".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
c94202ec — pdfinn 6 months ago
fix(llm): harden tool-use history to prevent cascading failures

- Record tool_results in history before the API call so orphaned
  tool_use blocks can't corrupt the session if the call fails
- Detect and auto-recover from tool_use/tool_result history mismatches
  by resetting the session and retrying
- Add synthetic assistant error message on AskWithToolResults failure
  to keep role-alternation valid
- Replace manual JSON escaping in buildToolResultsJSON with json.Marshal
- Update mock AskWithRequest signatures to return AskResponse

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
11a3967d — P. D. Finn KD9WEH 6 months ago
Merge pull request #1 from NERVsystems/claude/local-llm-feasibility-gVLhq

Add OpenAI-compatible local LLM backend support
2eac753a — Claude 6 months ago
Add OpenAI-compatible backend for local LLM support (GPT-OSS)

Implement OpenAIClient backend that speaks the OpenAI Chat Completions
API (/v1/chat/completions), enabling llm9p to work with any local model
server: Ollama, vLLM, llama-server, LocalAI, or LM Studio.

Primary target is GPT-OSS (OpenAI's open-weight MoE models), but any
model served via these platforms works. The backend supports:
- Blocking and streaming chat completions
- Tool/function calling with STOP:/TOOL: formatting
- Token counting from API usage (with estimation fallback)
- Conversation history, system prompts, temperature control
- Stateless AskWithRequest for session isolation

Usage:
  ./llm9p -backend openai -openai-url http://localhost:11434/v1 -model gpt-oss:20b

Also fixes pre-existing stale mock backends in test files (AskWithRequest
signature was out of date with the Backend interface).

https://claude.ai/code/session_017qVVZUUhfCCvNkMYmXDZAa
822b318b — pdfinn 6 months ago
fix(llmfs): clean up sessions on disconnect to prevent memory leak

Each spawned subagent creates an LLM session via /n/llm/new, but sessions
were never automatically freed: BaseFile.Close() was a no-op, and connection
drop handling only removed the client from the map without touching sessions.
Repeated spawn calls caused sessions to accumulate indefinitely, consuming
memory for full conversation histories.

Changes:
- Add refs int32 (atomic) to Session struct
- Add IncRef(id) / DecRef(id) to SessionManager; DecRef calls Close when
  refs reach zero, freeing the session and all its conversation history
- Add sessionRefFile / sessionRefDir wrappers (session_ref.go) that call
  IncRef on Open and DecRef on Close (once.Do guards against double-decrement)
- SessionsDir.Lookup wraps the returned SessionDir in sessionRefDir, so
  child file lookups via sessionRefDir.Lookup also yield sessionRefFile objects
- Fix handleConn disconnect path to call file.Close() on all unclosed fids
  before removing the client, ensuring sessions are freed on network drop

Typical flow: subagent exits → NEWFD closes all FDs → 9P connection drops →
handleConn deferred cleanup calls Close on remaining fids → sessionRefFile.Close
→ sm.DecRef → refs=0 → sm.Close(id) → session deleted from map.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
4191d4e7 — pdfinn 6 months ago
fix(llmfs): correct model alias IDs to match actual Anthropic API

Previous commit used claude-sonnet-4-6 and claude-opus-4-6 which don't
exist in the Anthropic API. IDs verified against anthropic-sdk-go v1.19.0:

  haiku  → claude-haiku-4-5-20251001    (unchanged, was correct)
  sonnet → claude-sonnet-4-5-20250929   (was claude-sonnet-4-6, 404)
  opus   → claude-opus-4-5-20251101     (was claude-opus-4-6, 404)

Also fix default from claude-sonnet-4-6 (invalid) to claude-sonnet-4-5-20250929.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
27f32368 — pdfinn 6 months ago
fix(llmfs): expand model aliases and update default model ID

Add a modelAliases map in session_settings.go so that short names
written to /n/llm/N/model are expanded to full Anthropic model IDs:

  haiku  → claude-haiku-4-5-20251001
  sonnet → claude-sonnet-4-6
  opus   → claude-opus-4-6

This fixes spawn subagents which write short names like "haiku" to the
model file — the Anthropic API rejects bare aliases as invalid model IDs.

Also update the default model from the stale claude-sonnet-4-20250514
to claude-sonnet-4-6 to match current API availability.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
bad7a7b7 — pdfinn 6 months ago
feat(llmfs): async Write + per-session streaming for live token delivery

Enable lucibridge to stream LLM tokens into the Lucifer conversation
zone as they are generated, without waiting for the full response.

Changes:
- session.go: add BeginGeneration/EndGeneration/SendChunk/GetStreamCh/
  WaitDone methods on Session; streamCh (cap 256) carries raw text chunks
  during generation; doneCh signals completion; EndGeneration closes
  channels but does NOT nil streamCh (late readers still see closed chan)
- session_ask.go: Write() is now async — calls BeginGeneration(), spawns
  goroutine, returns immediately; Read() calls WaitDone() before accessing
  LastResponse so pread blocks until generation completes
- session_stream.go (new): /n/llm/N/stream file; Read() blocks on <-ch
  returning each text chunk as it arrives; returns EOF when generation
  is done or no generation is active (channel nil or closed)
- session_dir.go: register stream file in Children() and Lookup()
- client.go: AskWithRequest() branches on req.StreamFunc != nil to use
  SSE Messages.NewStreaming() path; text_delta events forwarded to
  StreamFunc; session.Ask() sets StreamFunc=session.SendChunk when
  GetStreamCh() is non-nil
- server.go: add detailed walk debug logging (names, types, failures)
  controlled by existing -debug flag

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
0afa911a — pdfinn 6 months ago
feat(llmfs): native Anthropic tool_use protocol support

Add structured tool_use protocol to llm9p, enabling Veltro to use
Claude's native JSON tool invocation instead of text-based parsing.

New types (backend.go):
- AskResponse: carries Response, StructuredJSON, and Tokens — replaces
  the old (string, int, error) return from AskWithRequest
- ToolDef: tool definition passed to Anthropic tools API
- ToolResult: tool execution result for submission to the LLM
- Backend.AskWithRequest() now returns (AskResponse, error)

client.go:
- Message.StructuredContent: stores JSON content blocks for correct
  history replay of tool_use and tool_result turns
- AskWithRequest(): when ToolDefs non-nil, passes tools to API and
  returns STOP:/TOOL: formatted response for Limbo parsing
  Format: "STOP:tool_use\nTOOL:<id>:<name>:<args>\n<text>" or
          "STOP:end_turn\n<text>" or plain text (no tools)
- AskWithToolResults(): submits tool results as a new user turn
- Helpers: buildMessageParam(), buildToolParams(), extractToolArgs(),
  jsonEscapeString()

session.go:
- Session.tools field + SetTools/Tools methods
- Session.AddStructuredMessage() for storing structured content blocks
- AskRequest extended with ToolDefs and ToolResults fields
- SessionManager.Ask(): includes tools, stores structured JSON in history
- SessionManager.AskWithToolResults(): new method for tool result turns
- Fix Compact() for new AskResponse return type
- Helpers: extractTextContent(), buildToolResultsJSON()

cli_client.go: update AskWithRequest() to return AskResponse (no tools
support; StructuredJSON always empty)

session_tools.go (new): /n/llm/{id}/tools write-only file
- Write JSON array of ToolDef to enable native tool_use protocol
- Empty write clears tools (returns session to text-only mode)

session_ask.go:
- Detect TOOL_RESULTS\n prefix in Write() → parseToolResults() → AskWithToolResults()
- TOOL_RESULTS format: "TOOL_RESULTS\n<id>\n<content>\n---\n..."

session_dir.go: add tools file to Children() and Lookup()

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6da61fd4 — pdfinn 6 months ago
feat(llm9p): per-session compact and usage files for context window management

Add automatic context window compaction support to the per-session 9P API:

- session.go: Add Session.EstimatedContextTokens() (4 chars/token heuristic
  over current messages — more accurate than cumulative totalTokens for
  threshold decisions). Add SessionManager.Compact(ctx, id) which summarises
  the conversation via AskWithRequest then replaces session.messages with a
  compact 2-message exchange. Add SessionManager.EstimatedContextTokens(id)
  and SessionManager.ContextLimit() (200K for all Claude models).

- session_compact.go: New /n/llm/N/compact file. Write any content to
  trigger Compact() for that session. Follows the SessionModelFile pattern.

- session_usage.go: New /n/llm/N/usage file. Read returns
  "estimated_tokens/200000\n". Follows the SessionModelFile pattern.

- session_dir.go: Wire compact and usage into Children() and Lookup().

Also land two pre-existing uncommitted fixes:
- cli_client.go: Accept result messages with empty Result field
- protocol.go: Increase MaxMessageSize 8192→65536 for large system prompts

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
a3dc06aa — pdfinn 7 months ago
feat(llm9p): Implement clone-based session architecture

Replace per-fid session model with Plan 9 clone pattern:
- Reading /n/llm/new creates a session and returns its ID
- Each session gets its own directory: /n/llm/<id>/
- Per-session files: ask, ctl, model, system, thinking, context, metrics
- AskWithRequest method for stateless CSP-style LLM calls
- Session settings (model, temperature, thinking) are per-session
- Remove old ask.go, context.go in favor of session-scoped files

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ed43a61b — pdfinn 7 months ago
feat(llm9p): Add per-fid session isolation and prefill support

- Add SessionManager for per-fid conversation isolation
- Each 9P fid now gets its own conversation history
- Add FidAwareFile interface for files needing fid context
- Add /n/llm/prefill file for assistant response prefill
- Prefill helps keep model in character (e.g., "[Veltro]")
- Update ask, new, context files to use session manager
- Fix context contamination between parent and subagent

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>