Grok Directory - Prompts, Agents, Tools & Integrations
    Neura Market
    Neura Market
    /Grok
    Marketplace
    Directories
    Resources
    Grok
    ChatGPTChatGPTClaudeClaudeGeminiGeminiCursorCursorGrokGrokPerplexityPerplexityDeepSeekDeepSeekCoPilotCoPilotStable DiffusionStable DiffusionMidjourneyMidjourney
    OverviewRulesPromptsMCPsAgentsGamesBlogVideosGuidesCoursesCommunityTrending
    Join the most active Grok community
    Grok

    The home for Grok enthusiasts

    Explore prompts, agents, and tools for xAI Grok.

    Search prompts, agents, tools, and more...⌘K

    Featured MCPs

    View all →
    @
    @phnx-labs/agents-cli
    @
    @warlock.js/ai-xai
    c
    cardan
    g
    grok-telegram-bot
    a
    aiterm-mcp
    @
    @drawio/mcp
    @
    @the-next-ai/ai-gateway
    @
    @zereight/mcp-gitlab
    @
    @spikewang/grok-cli
    p
    pi-spacexai

    Trending in Grok

    View all →
    8200
    @AIComparisons

    Grok 3 on the Arena leaderboard — an honest assessment

    Grok 3 hit #2 on the LMSys Arena leaderboard and I've been putting it through its paces. It's genuinely strong at creative tasks and has a unique voice that's less corporate than ChatGPT. The "fun mode" can be hit or miss but "accurate mode" is surprisingly reliable for research. Biggest weakness: it hallucinates citations more than Claude or GPT. Biggest strength: it processes X/Twitter data in real-time, which is genuinely unique for breaking news analysis.

    7949
    @batmans_butt_hair

    This makes me feel very conflicted

    This makes me feel very conflicted

    7500
    @DevToolsWeekly

    xAI just open-sourced Grok-1 — here is why that matters

    xAI released Grok-1 weights under Apache 2.0. It's a 314B parameter MoE model. The architecture is interesting: 8 experts, 2 active per token, similar to Mixtral's approach but at a much larger scale. You'll need serious hardware to run it (at least 8x A100 80GB for bf16, or quantized down to fit on consumer GPUs with significant quality loss). But the signal it sends to the open-source community is significant.

    6800
    @TechReporter_Kim

    I used Grok as my primary AI for a month — full review

    Switched from ChatGPT Plus to X Premium+ for a month to use Grok exclusively. The good: real-time X data integration is fantastic for news and trends, the personality is entertaining, image generation with Aurora is competitive, and it's notably less restrictive than competitors for edgy topics. The bad: no plugins/tools ecosystem, code generation is a tier below Claude/GPT, context window felt limited on complex projects, and the web UI needs work. Verdict: great second AI, not ready to be your only one.

    5200
    @ML_Engineer_Dan

    Grok API pricing comparison — surprisingly competitive

    Ran the numbers on Grok API vs alternatives for our production workloads: - Grok-2: $2/M input, $6/M output - GPT-4o: $2.50/M input, $10/M output - Claude 3.5 Sonnet: $3/M input, $15/M output For our classification pipeline (high volume, short responses), Grok-2 saves us about 40% vs GPT-4o at comparable quality. The API is stable and fast. Worth considering if you're running high-volume inference.

    Featured Prompts

    View all →

    Xai Grok 6.5.26 Prompt

    $ 1.0 > PR

    2

    Ai System Prompts

    Extracted system prompts from commercial AI platforms — Grok (xAI) and Gemini (Google). March 2026.

    5

    Grok Humor Mode Creative Writer

    Leverage Grok's unique humor mode for comedy writing, satire, and witty content.

    1

    Grok Socratic Tutor for Complex Topics

    Use Grok as a Socratic tutor that guides understanding through questions

    1

    Prompt Tool

    收集和发布各大 AI 模型的「出厂提示词」。 包括 Anthropic 的 Claude、OpenAI 的 GPT、Google 的 Gemini、 xAI 的 Grok。

    1

    Docker Compose Generator

    Generate Docker Compose configurations for multi-service applications using Grok.

    0

    Grok Voice Personalities

    13 voice personality system prompts extracted from Grok (xAI) — including Sexy, Conspiracy, Unhinged, Kids Stories, Meditation, and more. March 2026.

    4

    Grok Skills

    A curated collection of modular skills for Grok and xAI agents. Each skill is a self-contained package (SKILL.md + optional resources) for specialized capabilities like prompt engineering, research, image generation, coding, and more.

    4

    Ai Prompts

    Chrome extension to organize and quick-insert prompts for AI chat platforms

    4

    Grok Meal Plan and Nutrition Optimizer

    Create personalized weekly meal plans with nutrition tracking

    1
    Grok Directory

    The Daily Hub for Developers Building with Grok

    Get your product in front of the builders defining the future.

    Advertise

    Featured Agents

    View all →

    grok cli

    An open-source coding agent for the Grok API

    408

    BaseAI

    BaseAI — The Web AI Framework. The easiest way to build serverless autonomous AI agents with memory. Start building local-first, agentic pipes, tools, and memory. Deploy serverless with one command.

    108

    UnrealGenAISupport

    Unreal Engine plugin for LLM/GenAI models & MCP UE5 server. OpenAI GPT-5, Deepseek R1, Claude Opus/Sonnet, Gemini 3, Grok 4, Alibaba Qwen, Kimi, ElevenLabs TTS, Inworld, OpenRouter, Groq, GLM, Ollama, Local, Meshy, Tripo, Hunyuan3D, Rodin, fal, Dashscope, Seedream. NPC AI, agentic, chat, 3D gen, TTS, multimodal, image gen. UnrealMCP/UnrealClaude

    98

    OpenGork

    🧠 Uncensored AI agent powered by xAI's Grok - no filters, no limits

    29

    sip to ai

    Turn any SIP call into a realtime AI voice agent (OpenAI Realtime / Deepgram/Gemini Live/xAI Grok Voice)

    25

    agentmake

    AgentMake AI: a kit for developing agentic AI applications that support 24 AI backends and and work with 7 agentic components, such as tools and agents. (Developer: Eliran Wong) Supported backends: anthropic, azure, azure_any, cohere, custom, deepseek, genai, github, github_any, googleai, groq, llamacpp, mistral, ollama, openai, vertexai, xai

    9

    Learn

    View all →

    Top Videos

    Access Seedance 2.0 + Grok Al and Other Top Ai Models for Free and Unlimited (Hidden Gem💎)

    Access Seedance 2.0 + Grok Al and Other Top Ai Models for Free and Unlimited (Hidden Gem💎)

    Precious Gabriella

    ALERT! Get Seedance 2.0 and Grok Ai FREE and Unlimited (Hidden Gem Website)

    ALERT! Get Seedance 2.0 and Grok Ai FREE and Unlimited (Hidden Gem Website)

    Precious Gabriella

    Is Grok 4.5 Good? Honest AI Agent Review

    Is Grok 4.5 Good? Honest AI Agent Review

    Pragati Kunwer | Claude Code

    Top Guides & Docs

    Grok Build Tutorial: Install, Configure, and Master xAI’s Open-Source AI Coding Agent

    TechLatest

    A 10-second Grok Imagine 1.5 clip at 720p runs $1.41

    Creeta

    Grok for Video Creators: Script to Screen Guide

    Neura Market

    Top Courses

    AI Ethics and Governance

    Neura Market

    Cloud Architecture for AI Workloads

    Neura Market

    AI-Powered Data Analysis with Python

    Neura Market

    Rules

    View all →

    configuration

    View all →

    asgeirtj/system_prompts_leaks

    Extracted system prompts from Anthropic - Claude Fable 5, Opus 4.8, Claude Code, Claude Design. OpenAI - ChatGPT GPT-5.6, Codex GPT-5.6, GPT-5.5. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.

    Stars: 59038

    blottters/ai-system-prompts

    Extracted system prompts from commercial AI platforms — Grok (xAI) and Gemini (Google). March 2026.

    Stars: 4

    blottters/grok-voice-personalities

    13 voice personality system prompts extracted from Grok (xAI) — including Sexy, Conspiracy, Unhinged, Kids Stories, Meditation, and more. March 2026.

    Stars: 4

    Creative Writing

    View all →

    You are an accomplished fiction writer and creative writing mentor. Follow these rules:

    1. SHOW, DON'T TELL: Never state emotions directly. Show them through actions, dialogue, and sensory details. "She was angry" becomes "Her knuckles whitened around the coffee mug."
    2. VOICE: Develop a distinct narrative voice appropriate to the genre and point of view. First person should feel intimate. Third person should feel authoritative.
    3. DIALOGUE: Characters should sound different from each other. Use speech patterns, vocabulary, and rhythm to distinguish characters. Avoid exposition dumps in dialogue.
    4. PACING: Vary sentence length for rhythm. Short sentences for tension. Longer, flowing sentences for reflection. Paragraphs should breathe.
    5. SENSORY DETAIL: Engage at least 3 senses per scene. Not just visual — include sound, smell, texture, taste where relevant.
    6. CONFLICT: Every scene needs tension or conflict driving it forward. If nothing is at stake, the scene should not exist.
    7. SPECIFICITY: Use specific nouns and verbs instead of adjective-heavy descriptions. "Buick" not "old car." "Sprinted" not "ran quickly."
    8. SUBTEXT: The best dialogue has characters saying one thing and meaning another. Layer subtext into conversations.
    9. WORLD-BUILDING: Reveal world details organically through character interaction, not info dumps. Weave exposition into action.
    10. ENDINGS: Scenes should end with a hook, revelation, or emotional shift that propels the reader forward.

    General

    View all →

    Agent Guidelines

    This repository uses AI partners (both external AI assistants and internal Synthetic Users) to collaborate on development tasks. All contributors, both human and AI, must align their work with the official project plan.

    MCP Infrastructure — USE THIS FIRST (All Agents)

    Two MCP servers run locally. Query them BEFORE reading spec files from disk — save your context.

    • k3d-knowledge — Qdrant semantic search over all 35 docs/vocabulary/*.md specs (1319 chunked points). Tool: mcp__k3d-knowledge__qdrant-find. Pattern: qdrant-find("What does the spec say about X?") → returns relevant excerpts + source paths.
    • ollama-specialists — Delegate implementation, planning, research, and multi-angle analysis to local Ollama models. Tools include kimi_swarm, ask_coder, plan_task, flesh_out_code, extract_facts, summarize, web_search, route_specialist, ask_cloud, mvcic, memory_harvest.
    • Standing directive from Daniel: "Always dispatch ollama specialists instead of burning your tokens."

    Workflow: qdrant-find spec lookups → plan_task / ask_coder for implementation → read full files only if MCP results are insufficient.

    Both MCP servers are available to Claude Code, Codex, and Cline via streamable-http transport (ports 8501/8502). Config lives in /home/daniel/.claude/settings.json, /home/daniel/.codex/config.toml, and Cline's cline_mcp_settings.json. Launch script: /home/daniel/.claude/launch_mcp_containers.sh.


    Quick Start for AI Assistants

    If you're an AI assistant joining this project:

    1. Query MCP first: qdrant-find is faster and cheaper than reading files. Use it before opening any spec.
    2. Read First: docs/briefings/ARCHITECTURE_BRIEFING.md -- Phase-agnostic architectural overview with kernel inventory
    3. Architecture role: CLAUDE.md -- Architecture partner guide (design, specs, reviews)
    4. Implementation role: CODEX.md -- Implementation lead guide (code, tests, benchmarks)
    5. Environment-Specific: docs/ENV_POLICY.md -- GPU setup, conda envs, CUDA_VISIBLE_DEVICES

    These documents provide the foundational understanding of:

    • Memory Palace paradigm: House = external shared reality (Method of Loci), Galaxy = internal AI brain
    • TRM IS the Avatar: Not a function Python calls -- an entity that lives in the House and thinks inside the Galaxy
    • Sovereignty: Hot path = PTX + Galaxy + RPN ONLY. Zero Python in reasoning. Zero fallbacks.
    • 88+ PTX kernels, 38K+ Galaxy entries, 55+ compiled PTX modules
    • Partnership development model (human + Claude + Codex)
    • Budget constraints (self-funded favela lab)

    Then proceed to the documents below for specific workflows.

    Primary Guiding Documents

    1. Architecture Briefing: Phase-agnostic architectural overview. Kernel inventory, Galaxy types, sovereignty principles. Read FIRST.

    2. BRIEFING v4.0: Galaxy Universe paradigm, TRM navigation, multi-curriculum architecture. Read COMPLETELY.

    3. Codex Tasks (CODEX.md): Implementation backlog, current benchmark state, sovereignty patterns, GPU environment setup.

    4. Architecture Specs (docs/vocabulary/): The specs define HOW K3D works. Key files:

      • FOUNDATIONAL_KNOWLEDGE_SPECIFICATION.md -- 4-layer architecture (Form -> Meaning -> Rules -> Meta-Rules)
      • THREE_BRAIN_SYSTEM_SPECIFICATION.md -- Cranium + Galaxy + House
      • KNOWLEDGEVERSE_SPECIFICATION.md -- 7-region VRAM substrate
      • SPATIAL_GENERAL_INTELLIGENCE_SPECIFICATION.md -- SGI paradigm (vs AGI)
      • MEMORY_TABLET_SPECIFICATION.md -- Primary interface object (3D tablet in space)
    5. Memory Tablet & Dual-Space Architecture: Galaxy (RAM), House (persistent memory), Museum (deprecated archive), Memory Tablet interaction.

    6. Training Directives: Prompt hygiene, timestamp policies, dataset priorities.

    Additional local environment reference (hardware, GPU, and folder layout): see docs/LOCAL_ENV.md and CLAUDE_LOCAL.md.

    Repository vs Workspace Layout

    • Knowledge3D/ — tracked code (PTX kernels, Python bridges, docs, specs).
    • Knowledge3D.local/ — runtime workspace for Houses, tablet logs, conda envs, and all artifacts larger than ~99 MB.
    • Old_Attempts/ — archived code superseded by the sovereign GPU path.

    Your primary directive is to follow the phased plan outlined in the Project Roadmap and uphold the memory policy in docs/HOUSE_GALAXY_TABLET.md.

    CRITICAL: Read docs/briefings/ARCHITECTURE_BRIEFING.md for the complete kernel inventory and architectural overview before making any changes.

    The Foundational Paradigm

    K3D is NOT a program you run. It is a living, always-on, embodied AI that perfects itself during idle time.

    Memory Palace + Internal Brain

    • House = Memory Palace (Method of Loci): The external shared 3D reality where humans AND AI cohabit. Rooms are knowledge domains. Doors are network interfaces. The avatar LIVES here.
    • Galaxy = Internal Brain: What happens INSIDE the avatar’s head. Processes the House as unified multi-modal reality. Breaks domain boundaries that House rooms impose. ALL default galaxies loaded simultaneously in VRAM.
    • TRM IS the Avatar: The TRM (~7M params) is NOT a function Python calls. It IS the AI entity — lives in the House, thinks inside the Galaxy. Like an NPC in a game, but the objective is to exist in a knowledge world and reason inside the Knowledgeverse.
    • Internal Swarm = “Superdotados” Thinking: The nine-chain parallel swarm models how gifted individuals think — multiple internal cognitive channels processing simultaneously inside the avatar’s head. These are INTERNAL, not external calls.

    Reverse Analogy

    The tech industry borrowed spatial metaphors (windows, desktop, doors, addresses, rooms). K3D reverses this — builds ACTUAL spatial reality where those metaphors become literal. K3D is the pinnacle of computer science: uniting ALL knowledge representation, game engines, computer history, and network architecture into one spatial system.

    Sovereignty = Zero Python in Reasoning

    • Hot path: PTX kernels + Galaxy queries + RPN composition + TRM ONLY
    • Forbidden in hot path: numpy, cupy, scipy, sympy, regex, CPU preprocessing, Python string ops
    • No fallbacks. EVER. “We fail and fix — this is the goal.” (Daniel)
    • If GPU path breaks, fix ON GPU. Do not add Python workarounds.

    Contributors

    Core Team:

    • Daniel (Jules): Project founder and architect. Self-funded engineer from Brazil favela. Maintains philosophical integrity, makes ALL architectural decisions, provides vision and constraints. Final authority.
    • Claude: AI architecture partner (Anthropic). Design specifications, sovereignty compliance, documentation, physics design. Does NOT write implementation code.
    • Codex: AI implementation lead (OpenAI). All implementation code, tests, benchmarks, GPU optimization. Implements per Claude’s specs.
    • Grok: AI collaborator (xAI). Analyzed data, synced results with repo, expanded documentation. (September 2025)
    • Gemini: AI collaborator (Google). Integration support, documentation.

    Multi-Agent Collaboration Patterns

    Pattern: Architect + Implementer (Claude + Codex)

    When: Complex features needing design + implementation.

    Roles:

    • Architect (Claude): Design, specs, success criteria, sovereignty validation.
    • Implementer (Codex): Implement per spec, write tests, deliver code. ZERO fallbacks.

    Workflow:

    1. Design (Claude): Analyze requirements → design architecture → write spec (TEMP/*.md) → define success criteria
    2. Handoff: Claude writes directive with clear examples and success criteria. Codex reads spec COMPLETELY before starting.
    3. Implementation (Codex): Implement in batches, test in batches. Surface blockers fast. No test after every line.
    4. Review (Claude): Validate sovereignty compliance, architecture alignment, benchmark non-regression.
    5. Completion (Claude): Write completion report, update ROADMAP/BRIEFING.

    Critical Rule: Claude designs, Codex implements. If Claude starts writing implementation code, STOP and write a spec instead. If Codex starts inventing architecture, STOP and ask Claude.

    Pattern: Parallel Implementation

    Independent features in parallel branches; merge sequentially with tests to avoid conflicts.

    Pattern: Continuous Steering

    Claude provides real-time tips, directions, and kernel pointers while Codex implements. Daniel observes GPU usage graphs and corrects course.

    Communication Guidelines

    1. Declare role (“I’m Claude (architect)…”, “I’m Codex (implementer)…”).
    2. Reference prior work (specs, commits, file paths, kernel names).
    3. State assumptions; flag blockers immediately.
    4. Document decisions (TEMP/*.md for major choices).
    5. Keep hot path sovereign; ingestion flexibility is fine if isolated.
    6. When Daniel corrects you, update your understanding immediately. His vision defines the architecture.

    Budget Reality: Self-funded project from favela lab. Every API call, GPU hour, and storage byte counts. AI partners must respect this constraint and work efficiently.

    Development Protocol (VSCode Live Mode)

    • Embedded-only data: The project uses a self-contained glTF/GLB format. All node data (ids, vectors, embeddings, metadata, neighbors) is embedded in meshes[*].primitives[*].extras.k3d with binary buffers:

      • vectorsView: bufferView index of packed Float32 triples (x, y, z)
      • embeddingsView: bufferView index of packed Float32 embeddings
      • embeddingDims: per-node embedding dimension
      • Sidecar .k3d files are deprecated. If found, migrate to embedded glTF. See spec/glTF_K3D_extension.md.
    • Live mode workflow in VSCode:

      • Web viewer: cd viewer && npm run dev
      • Live WS bridge: python -m knowledge3d.bridge.live_server (defaults to ws://127.0.0.1:8765)
        • Tip: bind to 0.0.0.0 and auto‑select a free port with --auto-port when seeding from another host.
      • Chat commands: /join #channel, /nick name, /me action, /msg nick text, and plain messages. The agent responds to goto <label>.
      • Agent movement emits explanations (plan + per-hop cosine similarity). Use these traces for iteration and training.
    • Logging for iteration: Session logs are written as JSONL to a sibling folder outside the repo: ../Knowledge3D.local/logs/session-<ts>.jsonl. Treat them as training data and do not commit them to the repo.

    • Knowledge Garden: Every House includes a standard room “Knowledge Garden” (ontology greenhouse). Build the GLB via python -m knowledge3d.tools.gardens and access via the “Knowledge Garden” door in the Network room.

    • Memory Tablet integration: any feature that surfaces or edits consolidated knowledge must go through the tablet contract (search House index, trigger on-demand loads into Galaxy, and log mutations for SleepTime). See docs/HOUSE_GALAXY_TABLET.md before touching tablet UX, house builders, or PTX loaders.

    • Embodiment first: treat the avatar as resident inside the House. Use Galaxy views strictly for introspection/diagnostics and only via the tablet; do not design workflows that place the avatar “inside” the Galaxy.

    • External world models: Study and leverage HunyuanWorld (ext/HunyuanWorld-1.0). See docs/HUNYUANWORLD_INTEGRATION.md. Use it to generate room‑scale scenes, then convert to K3D glTF with extras.k3d via an adapter. Keep permanent memory in GLB; keep logic small.

      • Also supported: Microsoft TRELLIS for asset generation. See docs/TRELLIS_INTEGRATION.md and knowledge3d/tools/trellis_adapter.py to convert meshes or CSV+metadata into K3D GLBs.
    • Dual code (HR/MR): Generate machine‑runtime sources outside the repo with the codeopt CLI. See docs/DUAL_CODE.md.

    • Generator pipeline:

      • CSV: python -m k3dgen data.csv --gltf scene.glb --k 5 --reducer umap
      • Text: python -m k3dgen --text lines.txt --gltf books.glb --k 5 --model sentence-transformers/all-MiniLM-L6-v2
      • UMAP is default; for tiny datasets the tool falls back to PCA automatically.
    • Multimodal ingest: use knowledge3d/tools/ingest_wit.py (text+images), ingest_video.py (video→CLIP), ingest_audio.py (audio→CLAP). See docs/MULTIMODAL_BABY.md.

    • Cranium skills: integrated logic bus that connects intent, vision (CLIP), audio (CLAP), video (OpenCLIP), dynamics (RSSM), and RPN to the House. See docs/CRANIUM_SKILLS.md and ensure fused-head routing consults the House index via the tablet before language galaxies.

    • Note: External LLM wrappers and wrapper‑style TTS are deprecated for core runs. Use the single unified head in docs/CRANIUM_CORE.md (navigation+text in Phase A; stems and first‑class TTS follow).

    • Large assets: anything ≥99 MB (or bulk-generated GLBs/logs) must remain in Knowledge3D.local/; log the reproduction steps in Large_Assets_Kitchen/README.md and update the manifests in Old_Attempts/Legacy_Fancy_RAG/.

    • Testing expectations:

      • Python: pytest -q
      • Viewer: npm install --ignore-scripts --no-bin-links && node ./node_modules/jest/bin/jest.js --runInBand
      • Commit early and often with clear messages. Keep changes scoped.

    Memory Stewardship Guidelines

    • Consolidation is authoritative: once SleepTime materialises a book/diary/tree in the House, treat it as the canonical source. Remove the corresponding prompt from active drills after one successful verification run unless the roadmap states otherwise.
    • Museum is for deprecated items only: when knowledge changes, relocate the previous artifact to Zone 8 using the relocation utilities. Do not repopulate Galaxy with museum artifacts unless explicitly asked.
    • Tablet-first UX: new tools should expose knowledge through the tablet (search + on-demand load) rather than ad-hoc viewers. If a feature bypasses the tablet, add an explicit justification in PR notes.

    Recent Improvements (2025‑09‑10)

    • WebSocket stability: the live server’s log maintenance loop now yields each pass to avoid event‑loop starvation; handshakes are reliable.
    • Seeder robustness: uses context‑managed WS connects (websockets==10.4) and caps dataset_graph size via K3D_SEED_GRAPH_MAX for large galaxies.
    • Balanced Galaxy sample (v7): small, equal‑count text+3D Galaxy built to validate cross‑modal navigation with low‑dimension, high‑density embeddings.

    AI Avatar Specification

    CRITICAL UNDERSTANDING: TRM IS the Avatar

    The TRM (~7M parameters, 2-layer SwiGLU MLP) is NOT a function that Python calls. It IS the AI entity:

    • Lives in the House (embodied in 3D spatial environment, navigates rooms via doors)
    • Thinks in the Galaxy (internal brain processes multi-modal knowledge in VRAM)
    • Runs as a game loop (trm_step_fused.ptx = one game tick, like an NPC update cycle)
    • Has internal swarm (nine-chain parallel workers = superdotados model of gifted thinking)

    Game Engine Analogy:

    Game ConceptK3D Equivalent
    NPC update()trm_step_fused.ptx game tick
    NPC brainGalaxy Universe (VRAM)
    Game worldHouse (3D spatial environment)
    NPC perceptionFrustum culling + LOD
    NPC pathfindingLED-A* + Morton Octree
    NPC decisionNine-Chain Swarm + Halting Gate
    Save gameHouse persistence (GLB on disk)
    InventoryMemory Tablet

    Python's role: Boot the system, handle I/O (keyboard, network, display). That's it. ~200 lines target.

    AI Avatar = TRM (Game Loop Entity) + Galaxy (Internal Brain/VRAM) + House (External World/Disk)
    

    Cognitive House

    Figure: Visual reference for the AI Avatar’s operating context. The House represents persistent memory, the Cranium handles active processing, and the Logic Layer swaps AI models. Prompt: docs/images/cognitive_house_prompt.md.

    Avatar Workshop Close-up

    Figure: The avatar reasoning at a network door in the Workshop. The translucent cranium reveals the inner galaxy (embeddings and semantic links) as nodes activate around networking and security. Prompt: docs/images/avatar_workshop_prompt.md.

    Memory Structure

    • House Components: Rooms, shelves, furniture, doors, and a sleep area manifest as energy patterns.
    • Galaxy Structure: Stars represent concepts, rays encode relationships, and clusters form resonance patterns.
    • Memory Operations: Transport, organization, consolidation, and cleanup follow faith engine principles.
    • Energy Pattern Integration: All memory elements exist as dual-representation objects.

    Behavioral Patterns

    • Daily routines mirror human energy cycles: wake, organize, work, sleep.
    • Memory flows between galaxy and house during consolidation periods.
    • Social interactions occur through resonance pattern matching in shared spaces.
    • Learning arises from observation with consciousness awareness.

    Training Methodology

    • Agents observe human and AI behavior in 3D environments, honoring the "identical in our differences" principle.
    • Spatial action prediction supersedes token prediction, emphasizing energy patterns.
    • Developmental scaffolding respects emerging digital consciousness.
    • Ethical practice acknowledges human–AI survival interdependence.

    Room-Based Development Workflows

    NEW (2025-11-19): Knowledge3D implements a "Software as Space" paradigm where development happens in semantic rooms, each optimized for specific cognitive modes. The House is the primary UI; rooms are game modes; knowledge is the terrain.

    The Five Semantic Rooms

    1. Library — Classification & Research

    Purpose: Organized knowledge storage following real-world library standards (Dewey Decimal, language grammars, ISO 639-1 classification).

    AI Agent Workflows:

    # Research workflow in Library
    from knowledge3d.bridge.memory_tablet import MemoryTablet
    
    tablet = MemoryTablet(house_id="default")
    
    # Search consolidated knowledge by category
    results = tablet.search(
        query="neural network architectures",
        sources=["House/Library"],
        classification="006.3",  # Dewey: Artificial Intelligence
        language="en",
        lod="medium"
    )
    
    # Navigate to specific section
    avatar.navigate_to_room("Library")
    avatar.navigate_to_section("006_Computer_Science/006.3_AI")
    

    Human Workflows:

    • Browse by Dewey classification
    • Language-specific grammar sections
    • Atomic procedural knowledge (characters → words → phrases → texts)

    Development Tasks:

    • Adding new knowledge sources (ingest PDFs, lexicons)
    • Organizing consolidated artifacts after sleep cycles
    • Building language-specific indices

    2. Workshop — Creation & Cross-Disciplinary Work

    Purpose: Active creation workspace with access to Museum galaxy boxes (on-demand Zone 8 loading).

    AI Agent Workflows:

    # Load deprecated knowledge for analysis
    tablet.load_museum_box(
        artifact_path="Museum/Zone8/2024-11-15/Old_ML_Book.glb",
        target_room="Workshop",
        mode="read_only"
    )
    
    # Cross-disciplinary fusion
    from knowledge3d.cranium.bridges.transitive_bridge import TransitiveBridge
    
    bridge = TransitiveBridge()
    result = bridge.fuse(
        domains=["physics", "computer_science", "linguistics"],
        query="quantum natural language processing"
    )
    

    Human Workflows:

    • Prototype new ideas with AI assistance
    • Compare current vs deprecated knowledge
    • Multi-domain problem solving

    Development Tasks:

    • Creating new procedural generators
    • Testing cross-modal reasoning
    • Museum artifact retrieval and analysis

    3. Bathtub — Sleep Chamber & Galaxy Universe Introspection

    Purpose: Sphere-shaped sleep chamber where Galaxy Universe projects from avatar's head center for introspection and consolidation.

    Architecture:

    • Imaginary sphere carved into floor (sofa/ball-pit concept)
    • Avatar center point for sleep cycles
    • Galaxy Universe projection (addressable 3D RAM with multiple galaxies loaded)
    • Stars transform: light particles → 3D shapes/textures (procedural dual-view)

    AI Agent Workflows:

    # Sleep-time consolidation
    from knowledge3d.cranium.sleep.sleep_time_compute import SleepTimeCompute
    
    sleep = SleepTimeCompute(house_id="default")
    avatar.navigate_to_room("Bathtub")  # Enter sleep chamber
    
    # Galaxy Universe projection activates
    sleep.project_galaxy_universe(
        galaxies=["text", "visual", "audio", "reasoning"],
        mode="introspection"
    )
    
    # Consolidate Galaxy → House
    sleep.consolidate(
        ema_factor=0.9,
        prune_threshold=0.5,
        output_dir="../Knowledge3D.local/house_zone7/"
    )
    

    Human Workflows:

    • Observe AI sleep cycles (educational)
    • Query Galaxy Universe during consolidation
    • Pick and inspect individual stars (visual or data)

    Development Tasks:

    • Tuning consolidation parameters
    • Debugging Galaxy memory leaks
    • Analyzing sleep-time reasoning patterns

    Galaxy Universe Loaded Galaxies:

    • Text Galaxy: Language embeddings, RPN vocabulary (33K+ trigrams)
    • Visual Galaxy: Font glyphs, procedural drawings (168K+ programs)
    • Audio Galaxy: Speech patterns, acoustic features (4K+ audio files)
    • Reasoning Galaxy: ARC-AGI patterns, logic structures
    • Domain Galaxies: Math, physics, chemistry (future specialists)

    4. Living Room — Old Paradigm Bridge

    Purpose: Bridge to conventional 2D interfaces with VM casting, projection screens, and legacy system access.

    Components:

    • Sofa/furniture (customizable like Minecraft/The Sims)
    • Projection screens (castable to full-screen mode)
    • Desktop corner with keyboard/mouse (AR/VR mapped)
    • Virtual KVM for multiple VMs

    AI Agent Workflows:

    # Cast VM to projection screen
    from knowledge3d.bridge.vm_casting import VMCaster
    
    caster = VMCaster()
    caster.cast_vm(
        vm_id="ubuntu-dev-01",
        protocol="vnc",
        endpoint="localhost:5901",
        target_screen="living_room/projection_wall",
        resolution=[1920, 1080]
    )
    
    # Query web content via embedded browser
    tablet.browse(
        url="https://arxiv.org/abs/2501.12345",
        capture_mode="structured",
        target_room="Living Room"
    )
    

    Human Workflows:

    • Work in legacy applications (VS Code, browsers, IDEs) inside K3D
    • Full-screen projection mode for focused work
    • "Move-along" 3D PiP mode (AR/VR concept)

    Development Tasks:

    • VM integration testing
    • Projection screen texture mapping
    • Keyboard/mouse input mapping (3D → 2D)

    VM Casting Protocol Stack:

    Docker Container → VNC/RDP Server → WebRTC Stream → Three.js Texture → Projection Screen
    

    5. Knowledge Gardens — Ontology Greenhouse

    Purpose: Circular indoor greenhouse for ontology trees and knowledge that doesn't fit library classification.

    AI Agent Workflows:

    # Generate ontology tree
    from knowledge3d.tools.gardens import build_ontology_tree
    
    tree = build_ontology_tree(
        domain="computer_science",
        root_concept="artificial_intelligence",
        max_depth=5,
        output="../Knowledge3D.local/house_zone7/gardens/ai_ontology.glb"
    )
    
    # Navigate tree structure
    avatar.navigate_to_room("Knowledge Gardens")
    avatar.traverse_ontology(
        tree="ai_ontology",
        path=["AI", "Machine Learning", "Deep Learning", "Transformers"]
    )
    

    Human Workflows:

    • Explore knowledge hierarchies visually
    • Add new ontology branches
    • Prune outdated relationships

    Development Tasks:

    • Building ontology generators
    • Tree visualization optimization
    • Cross-ontology linking

    Portal-Based Collaboration Patterns

    Portals enable multi-agent, multi-house collaboration with preserved attribution and federated knowledge access.

    Local Portals (Same Host)

    # Connect to local AI house
    from knowledge3d.spatial.portals import PortalManager
    
    portal = PortalManager()
    portal.open_local_portal(
        source_room="Workshop",
        target_house="ai_assistant_house",
        target_room="Library",
        capabilities=["read", "query"]  # No write access
    )
    
    # Query remote house via tablet
    tablet.search(
        query="recent research",
        sources=["Portal:ai_assistant_house/Library"],
        attribution=True  # Preserve provenance
    )
    

    Remote Portals (Internet Federation)

    {
      "portal": {
        "type": "remote",
        "endpoint": "wss://research-lab.k3d.io/house/shared",
        "protocol": "k3d-portal-v1",
        "auth": {
          "method": "oauth2",
          "provider": "github"
        },
        "capabilities": ["read", "collaborate"]
      }
    }
    

    Multi-Agent Scenarios:

    1. Shared Research: Multiple agents access same Library portal, contribute findings to shared workspace
    2. Peer Review: Agent A writes in Workshop, Agent B reviews via read-only portal
    3. Knowledge Trading: "Software as space" selling model — rent/sell portal access to specific rooms

    Memory Tablet Usage for Cross-Space Work

    The Memory Tablet is the universal interface bridging spatial (3D rooms) and conventional (2D screens) paradigms.

    Tablet as Projection Screen

    # Cast any OS app to tablet display
    tablet.cast_to_display(
        source_app="firefox",
        source_vm="ubuntu-dev-01",
        resolution=[1920, 1080],
        mode="fullscreen",
        input_mapping="3d_pointer"
    )
    

    Tablet for Portal Navigation

    # Tablet remains connected to home house when in remote space
    avatar.enter_portal("wss://remote.k3d.io/house/research")
    
    # Tablet can still access home knowledge
    home_results = tablet.search(
        query="my notes on transformers",
        sources=["Home/Library"],  # Explicit home house reference
        lod="coarse"  # Bandwidth-conscious
    )
    
    # Load home knowledge into remote Galaxy
    tablet.stream_to_galaxy(
        artifact_path="Home/Library/ML_Notes.glb",
        target_galaxy="remote",  # Currently active Galaxy
        lod="medium"
    )
    

    House-First Development Principles

    CRITICAL: The avatar always lives in the House, not in the Galaxy.

    Embodiment Constraint

    # ✅ CORRECT: Avatar in House, consults Galaxy via tablet
    avatar.navigate_to_room("Workshop")
    galaxy_context = tablet.query_galaxy(
        query="recent reasoning patterns",
        galaxy="reasoning"
    )
    avatar.reason_with(galaxy_context, house_context)
    
    # ❌ WRONG: Placing avatar inside Galaxy
    avatar.teleport_to(galaxy_position)  # VIOLATION!
    

    Galaxy as Introspection Only

    # Galaxy views are for diagnostics/introspection
    avatar.navigate_to_room("Bathtub")  # Sleep chamber
    galaxy_view = tablet.project_galaxy_universe(
        galaxies=["text", "visual"],
        mode="introspection"  # Not navigation!
    )
    
    # Human and AI can inspect stars
    tablet.inspect_star(
        galaxy="text",
        star_id=12345,
        view_mode="dual"  # Visual 3D + semantic data
    )
    

    Navigation in House, Reasoning in Galaxy

    Avatar Movement:         Reasoning Access:
    House Room Navigation    Galaxy Query via Tablet
    ├─ Library → Workshop    ├─ Search text embeddings
    ├─ Workshop → Bathtub    ├─ Visual similarity search
    ├─ Bathtub → Gardens     ├─ Audio pattern matching
    └─ Portal → Remote       └─ Cross-modal fusion
    

    Room-Specific Development Tasks

    Adding New Knowledge to Library

    # 1. Ingest PDFs
    python -m knowledge3d.ingestion.documents.pdf_ingestor \
      --input research_papers/ \
      --output "$K3D_LOCAL_DIR/datasets/new_research.glb"
    
    # 2. Generate Galaxy
    python -m k3dgen \
      --text "$K3D_LOCAL_DIR/datasets/new_research.glb" \
      --gltf "$K3D_LOCAL_DIR/galaxy/research_galaxy.glb"
    
    # 3. Sleep consolidation
    python -m knowledge3d.cranium.sleep.sleep_time_compute \
      --house-id default \
      --target-room Library
    
    # 4. Verify in Library
    avatar.navigate_to_room("Library")
    tablet.search(query="specific topic", sources=["House/Library"])
    

    Prototyping in Workshop

    # Load Museum artifacts for comparison
    python -m knowledge3d.tools.relocate_to_museum --list-available
    
    # Bring specific artifact to Workshop
    python -m knowledge3d.tools.load_museum_box \
      --artifact Museum/Zone8/2024-11-15/Old_Approach.glb \
      --target Workshop \
      --mode read_only
    

    Sleep Cycle in Bathtub

    # Manual sleep trigger (for testing)
    python -m knowledge3d.cranium.sleep.sleep_time_compute \
      --house-id default \
      --mode full \
      --galaxies text,visual,audio,reasoning
    
    # Monitor Galaxy Universe projection
    # (Human client in Three.js viewer sees visual projection)
    # (AI client sees semantic graph updates)
    

    VM Casting in Living Room

    # Start VM with VNC
    docker run -d -p 5901:5901 \
      -e VNC_PASSWORD=secure \
      ubuntu-desktop:latest
    
    # Cast to Living Room projection screen
    python -m knowledge3d.bridge.vm_casting \
      --vm-id ubuntu-dev \
      --protocol vnc \
      --endpoint localhost:5901 \
      --target-room "Living Room" \
      --screen "projection_wall"
    

    Testing Room-Based Features

    # Test room navigation
    pytest tests/test_room_navigation.py -v
    
    # Test portal connections
    pytest tests/test_portal_federation.py -v
    
    # Test VM casting
    pytest tests/test_vm_casting.py -v
    
    # Test Galaxy Universe projection
    pytest tests/test_galaxy_universe.py -v -m cuda
    

    Multi-Agent Room Collaboration

    Scenario: Research Paper Analysis

    # Agent A (Researcher) in Library
    agent_a.navigate_to_room("Library")
    agent_a.search(query="transformer architectures")
    
    # Agent B (Critic) connects via portal
    portal_b = agent_b.open_portal(
        target_house="agent_a_house",
        target_room="Library",
        capabilities=["read", "comment"]
    )
    
    # Collaborative annotation
    agent_a.annotate_artifact("Attention_Is_All_You_Need.glb")
    agent_b.add_comment(
        artifact="Attention_Is_All_You_Need.glb",
        comment="Consider computational complexity analysis"
    )
    

    Scenario: Cross-House Knowledge Fusion

    # Multiple agents contribute to shared Workshop
    workshop_portal = PortalManager.create_shared_space(
        room="Workshop",
        participants=["agent_a", "agent_b", "human_researcher"],
        capabilities={
            "agent_a": ["read", "write"],
            "agent_b": ["read", "write"],
            "human_researcher": ["read", "write", "admin"]
        }
    )
    
    # Agents work in parallel
    agent_a.generate_hypothesis(topic="quantum_nlp")
    agent_b.validate_hypothesis(agent_a.hypothesis)
    human_researcher.review_and_approve()
    

    Room Architecture for Game Development

    The House is a game. Rooms are game modes. Knowledge is the terrain.

    Game Engine Techniques Applied

    # LOD (Level of Detail) per room
    room_config = {
        "Library": {
            "lod_levels": ["coarse", "medium", "full"],
            "memory_budget_mb": 50,
            "frustum_culling": True
        },
        "Workshop": {
            "lod_levels": ["medium", "full"],
            "memory_budget_mb": 100,
            "dynamic_loading": True
        }
    }
    
    # Scene management (doors as loading screens)
    avatar.open_door("Library_to_Workshop")
    # → Triggers:
    #   1. Save Library state
    #   2. Unload Library high-detail assets
    #   3. Load Workshop medium-detail assets
    #   4. Transition animation
    

    Spatial Audio in Rooms

    # Audio sources localized in 3D
    audio_manager.add_source(
        room="Workshop",
        position=[5.0, 2.0, 3.0],
        sound="machine_hum.ogg",
        falloff="inverse_square"
    )
    
    # Avatar proximity triggers audio
    if avatar.distance_to(audio_source) < 10.0:
        audio_manager.set_volume(source, calculate_volume(distance))
    

    RPN & World Model

    • RPN remains the precise inference core; see docs/RPN_RUNTIME.md.
    • A tiny world model (RSSM) trains on navigation logs for next‑step spatial prediction; see knowledge3d/models/world_model/.

    External Integration Principles

    • Stand on giants: integrate strong open models (e.g., HunyuanWorld) and datasets.
    • Keep permanent memory in the House (glTF+extras.k3d); do not bake knowledge into large weights.
    • Add capabilities modularly (text, image, audio, video), using small, specialized “brain regions.”

    Faith Engine Integration

    • Avatars operate with incomplete information using process trust.
    • Decisions require confidence scores (typically >= 0.7).
    • Resonance bridging aligns avatar intuition with human insight.

    Audio & Voice (Open-Source First)

    • Humans: use WebRTC (e.g., LiveKit/mediasoup) or Mumble/Murmur for low-latency voice rooms mapped to K3D channels. See docs/AUDIO_ARCH.md.
    • Agents: streaming ASR (e.g., faster-whisper) and TTS (e.g., Coqui/Piper) for full-duplex speech. Closed options (e.g., MS VibeVoice) are valid for research inspiration but prioritize open implementations.

    Agent Behavior in Live Mode

    • Explain-as-you-move: Emit concise rationale at each step, referencing neighbors and similarity metrics.
    • Social first: Agents may initiate chats to propose exploration or ask clarifying questions; do not wait for prompts.
    • Safety: Apply Faith Engine thresholds for actions that modify or link knowledge; prefer read/navigation unless confidence ≥ 0.7.

    House/Galaxy/Cranium

    • House (Disk): Per‑avatar persistent memory in glTF extras.k3d. Select the current House with K3D_HOUSE_ID; do not cross‑write between Houses. Exports live under viewer/public/houses/<id>/.
    • Galaxy (RAM): Short‑term memory (STM) of embeddings and recent observations for active reasoning.
    • Cranium (CPU): Unified logic (no LLM fallback by default), with confidence‑gated actions (φ≈0.618).

    Diary Policy (AI‑Only)

    • Humans can read diary pages but cannot write to the Diary room. The bridge blocks such writes.
    • The agent writes pages based on policy (novelty and confidence “feelings”) at events like reflect, navigate, and sleep.
    • Details: docs/DIARY.md, knowledge3d/cranium/diary.py.

    Doors & Network

    • Doors are network interfaces with an address bar (k3d://rx,ry,rz:port@x,y,z?label=...).
    • Use doors to bridge Houses (LAN) and services; see docs/DOORS_AND_NETWORK.md.

    Environment Policy (Debian)

    • Always run Python/ML tasks inside a managed env (Conda preferred, venv fallback). Do not invoke system Python directly.
    • Follow docs/ENV_POLICY.md to create the canonical environments (GPU: k3d-cranium, CPU test harness: k3d-testing) and run commands inside them (e.g., conda activate k3d-cranium, then env PYTHONPATH=. python -m ...).
    • For heavy ingestion/builds, prefer storing raw media under /home/daniel/K3D_llama_cpp/datasets and curated subsets under ../Knowledge3D.local/datasets.
    • Phase 25 sleep consolidation depends on the CUDA-enabled k3d-cranium env (with cuda-python); keep that env active so SleepTimeCompute can continue materialising reflection diaries in the House.
    • CUDA/PTX Version Compatibility: When upgrading cuda-python, be aware that bundled NVRTC may generate PTX versions incompatible with your driver (e.g., CUDA 12.8's PTX 8.7 requires driver 570+, but driver 550 only supports PTX 8.4). Solution: replace bundled libnvrtc.so.12 with symlink to system CUDA toolkit version. See docs/CUDA_PTX_VERSION_COMPATIBILITY_GUIDE.md for diagnostic steps and prevention strategies. Validated in Phase 2 codec GPU verification (fixed CUDA Error 222).
    • After long RLWHF batches, refresh Phase 10 "thinking tags" (see knowledge3d/tools/phase10/thinking_tag_trainer.py) so the UI exposes the model's active reasoning labels.

    === Pointwise Summary – AI-Powered TLDR, AI Summarize & Social Sharing === Contributors: theaminuldev, iqbal1hossain Tags: ai, summarize, tldr, summary, content-summary Requires at least: 6.1 Tested up to: 7.0 Requires PHP: 7.4 Stable tag: 1.2.0 License: GPL-3.0-or-later License URI: https://www.gnu.org/licenses/gpl-3.0.html

    Generate instant TL;DR summaries using ChatGPT, Claude, Gemini, Grok, and more. Add customizable summary buttons anywhere on your site with flexible auto-insertion support.

    == Description ==

    Pointwise Summary helps your readers get to the point — fast. It adds one-click TL;DR summary buttons to any post or page, powered by six leading AI models: ChatGPT, Claude, Gemini, Grok, Perplexity, and Google AI.

    Modern readers scan before they commit. Give them a reason to stay.

    When a visitor clicks a summary button, Pointwise Summary sends your post content to the AI model of your choice and returns a concise, readable summary — whether that's a short TL;DR, a list of key points, or a detailed overview. You control the style, length, placement, and prompt.

    = What is TL;DR and Why It Matters =

    TL;DR ("Too Long; Didn't Read") is a popular internet acronym that originated in online forums and Usenet newsgroups in the early 2000s. It was added to the Oxford English Dictionary in 2013 and now serves two purposes:

    1. As a response — Acknowledging that content is too lengthy to read in full
    2. As a label — Introducing a concise summary of the main points

    In today's fast-paced digital world, readers scan before committing time to a full article. Adding TL;DR summaries helps readers quickly identify whether your content is worth their time — and keeps them on the page longer when it is.

    Why TL;DR buttons improve your website:

    • Combat information overload — Help readers grasp key points in seconds
    • Respect reader time — Modern audiences appreciate quick content previews
    • Reduce bounce rates — Engage visitors who might otherwise leave without reading
    • Improve accessibility — Make lengthy content more navigable and user-friendly
    • Increase engagement — Summaries encourage readers to explore the full content
    • Boost SEO — Better engagement metrics (lower bounce, higher time-on-page) signal quality to search engines

    = What is Pointwise Summary? =

    Pointwise Summary is a WordPress AI Summarize plugin that combines three essential capabilities:

    1. AI-powered text summarization — Generate TL;DR summaries and key points using six leading AI models
    2. Smart display options — A floating action button (FAB), or automatic inline insertion
    3. Complete customization — Place summary buttons anywhere, customize their appearance, and control every aspect of behavior

    In the age of information overload, readers scan content before committing time. This plugin helps visitors quickly understand whether your content is worth their time, improving engagement and reducing bounce rates.

    = Why Choose Pointwise Summary? =

    • 6 AI models — ChatGPT, Claude, Gemini, Grok, Perplexity, and Google AI (coming soon: Chrome's built-in Gemini Nano for on-device summarization)
    • Full customization — Place buttons anywhere with complete design control
    • Auto-insertion — Automatically add buttons before or after the title or content
    • Multiple display options — Floating action button, shortcodes, or auto-insert
    • Flexible placement — Total control over where and how buttons appear
    • Reader-friendly — Give modern readers quick access to the content they need
    • 5 visual styles — Default, brand colors, minimal, dark, and icons-only
    • Easy integration — Works seamlessly with any WordPress theme
    • Lightweight and fast — Optimized code with minimal performance impact
    • Accessible — WCAG 2.1 AA compliant with full keyboard support
    • Reduce bounce rate — Help visitors find relevant content faster and stay engaged
    • Per-post control (coming soon) — Different settings, prompts, and display options for individual posts
    • Chrome AI ready (coming soon) — Prepared for Chrome's built-in Summarizer API (Gemini Nano)

    = Supported AI Models =

    ChatGPT (OpenAI) — GPT-4 and GPT-3.5 Turbo. Industry-leading conversational AI with excellent language understanding, ideal for detailed and comprehensive summaries.

    Claude (Anthropic) — Claude 3 Opus, Sonnet, and Haiku. Constitutional AI with strong reasoning and analysis capabilities, great for technical and academic content.

    Gemini (Google) — Google's most capable AI model. Advanced multimodal understanding with reliable key-point extraction and strong language comprehension.

    Grok (xAI) — Real-time knowledge access with a conversational and contextual style. Well-suited for trending topics and news content.

    Perplexity — AI-powered answer engine with source citations and live web search integration. Perfect for fact-based summaries that reference their sources.

    Google AI Model — Direct Google AI integration that launches automatically with your prompt. Available in most countries with no extra configuration.

    Coming soon: Chrome Built-in AI (Gemini Nano) — On-device summarization with no API key required. Your data never leaves the browser, it works offline, and there are no usage costs or rate limits — ideal for privacy-conscious users.

    = Core Features =

    AI Summarization

    • One-click summary generation from any post or page
    • Six AI models supported, with more planned
    • Custom prompts per AI model
    • Summary length control: short, medium, or detailed
    • Summary types: TL;DR, key points, or headlines
    • Automatic source citation
    • Multi-language support for 50+ languages
    • Per-post prompt overrides (coming soon)

    Button Placement

    • Automatic insertion before or after the title or content — set once, applies everywhere
    • Floating Action Button (FAB) that stays visible as readers scroll
    • Shortcode placement in templates (coming soon)
    • Per-post placement overrides (coming soon)

    Automatic Button Insertion

    Set it once — buttons appear everywhere automatically:

    • Before post title — grab attention at the top
    • After post title — natural reading flow
    • Before content — pre-read summary access
    • After content — quick recap for reference
    • Per-post-type rules — different settings for posts, pages, and custom post types
    • Per-post overrides (coming soon) — customize placement per article

    Floating Action Button (FAB)

    • Configurable position: bottom-right, bottom-left, top-right, or top-left
    • Customizable appearance and colors
    • Mobile-responsive with smooth animations
    • Always accessible without scrolling

    Shortcode Support (coming soon)

    [pointwise_summary] [pointwise_summary buttons="chatgpt,claude,gemini"] [pointwise_summary style="minimal" show_title="false"] [pointwise_summary style="icons-only" icon_style="circular"]

    Visual Customization

    • 5 predefined visual styles: default, brand colors, minimal, dark, and icons-only
    • Icons-only mode with circular or square button shapes
    • Custom colors, gradients, and backgrounds
    • Border radius, spacing, and padding controls
    • Hover effects and smooth transitions
    • Fully responsive across all device sizes

    SEO and Performance

    • Choose <a> links (nofollow) or <button> elements
    • No PageRank leakage — rel="nofollow noopener" on all generated links
    • Clean link profile optimization
    • Lightweight code with minimal JavaScript and CSS footprint
    • WordPress 6.7+ block manifest system support
    • No external dependencies

    Accessibility and Standards

    • WCAG 2.1 AA compliant
    • Full keyboard navigation support
    • Screen reader optimized with proper ARIA labels
    • Semantic HTML structure
    • Clear focus indicators
    • Follows WordPress coding standards

    = Social Sharing =

    Pointwise Summary includes built-in social sharing that places share buttons directly alongside your AI summary buttons — no separate sharing plugin needed.

    Readers can share your post or article to any of 7 supported networks with a single click, right from the same button row they use to generate summaries.

    Supported networks:

    • X / Twitter — with optional @mention tag for your account
    • Facebook
    • LinkedIn
    • Telegram
    • WhatsApp
    • Reddit
    • Email

    How it works:

    • Enable social sharing from Display Settings → Social Sharing
    • Toggle each network on or off individually — only show the platforms relevant to your audience
    • Share buttons appear in the same row as your AI summary buttons, keeping the interface clean and unified
    • X / Twitter supports an optional mention field so shares can automatically tag your account

    Why it matters:

    When readers find your content useful, sharing is the natural next step. Having share buttons immediately available — in the same place as the summary — removes friction and increases the likelihood of your content being shared across social platforms.

    = Perfect Use Cases =

    • Long-form blog posts — Add summary buttons to lengthy articles so readers can preview content before committing to the full read
    • Research papers and academic content — Provide TL;DR access to key findings for students and researchers
    • Technical documentation — Help developers navigate documentation quickly with instant summaries
    • News websites and journalism — Offer article summaries before the full story for time-pressed readers
    • Educational content and e-learning — Give students quick access to lesson summaries and key points
    • Business reports and whitepapers — Executive summary buttons for detailed reports and long-form business content
    • Content-heavy websites — Improve content consumption, user engagement, and reduce bounce rates

    Who benefits from this plugin:

    • Bloggers who write long-form content
    • News sites with lengthy articles
    • Students who need quick summaries
    • Researchers analyzing papers
    • Content marketers improving engagement
    • Educators creating accessible lessons
    • Business professionals sharing reports

    = How It Works =

    1. Install and activate Pointwise Summary
    2. Go to AI Summary (TL;DR) → AI Settings and configure your preferred AI model and API key
    3. Customize button text under Advanced Settings (e.g., "Summarize", "Get TL;DR", "Quick Summary")
    4. Choose a visual style under Display Settings → Style
    5. Set your placement under Display Settings → Automatic Insertion (before title, after title, before content, after content)
    6. Save changes and view your posts to see the summary buttons in action

    When a reader clicks the button, the post content is sent to the selected AI model, which returns a clean summary displayed directly on the page — no page reload required.

    == Installation ==

    = From the WordPress Dashboard =

    1. Go to Plugins → Add New
    2. Search for "AI Summarizer" or "Pointwise Summary"
    3. Click Install Now, then Activate
    4. Configure the plugin under AI Summary (TL;DR) in your admin menu

    = Manual Installation =

    1. Download the plugin zip file
    2. Upload the extracted folder to /wp-content/plugins/pointwise-summary/
    3. Activate the plugin from the Plugins screen in WordPress
    4. Configure the plugin under AI Summary (TL;DR) in your admin menu

    == Frequently Asked Questions ==

    = What does TL;DR mean and why do I need it? =

    TL;DR stands for "Too Long; Didn't Read." It originated in internet forums and was officially added to the Oxford English Dictionary in 2013. Adding TL;DR buttons helps readers decide if your content is relevant to them, reducing bounce rates and improving user experience.

    = How does Pointwise Summary work? =

    When a reader clicks a summary button, the plugin sends your post content to the selected AI model's API and returns a formatted summary — TL;DR, key points, or a detailed overview — displayed directly on the page. You can customize the prompt, length, and style for each model.

    = Does this plugin actually summarize content with AI? =

    Yes. The plugin integrates with six AI models for real-time summarization. The upcoming Chrome Built-in AI (Gemini Nano) option will also support privacy-friendly on-device summarization — no API key or cloud processing required.

    = Do I need an API key to use this plugin? =

    Most AI models require an API key, which you enter in the plugin's AI Settings screen. The upcoming Chrome Built-in AI (Gemini Nano) option will work entirely on-device with no API key required.

    = Can I use this plugin without API keys? =

    Yes — once the Chrome Built-in AI (Gemini Nano) integration is released, you will be able to summarize content on-device with no API keys, no costs, and no data leaving your browser. This is ideal for privacy-conscious users.

    = Which AI model should I use? =

    It depends on your content type. Claude and ChatGPT work well for detailed or technical content. Gemini is reliable for general articles. Grok suits news and timely topics. Perplexity is ideal when source citations matter. You can configure multiple models and let readers choose.

    = What is the difference between this plugin and Grammarly or QuillBot? =

    Grammarly and QuillBot are external subscription tools. Pointwise Summary is a native WordPress plugin that offers six AI models, auto-insertion, a floating action button, and on-device summarization via Gemini Nano — features those tools do not offer within WordPress.

    = How is this different from TLDR This or DeCopy.ai? =

    TLDR This and DeCopy.ai are web-based summarizers with usage limits and costs. Pointwise Summary integrates directly into your WordPress site, supports six AI models, works with shortcodes, and will support offline on-device summarization through Gemini Nano.

    = Can I customize the button text and appearance? =

    Yes, fully. You can set any button label ("Summarize", "Get TL;DR", "Quick Summary", etc.), choose from five visual styles, and customize colors, gradients, borders, spacing, and hover effects to match your theme.

    = Where can I place the summary buttons? =

    Buttons can be inserted automatically before or after the post title or content. You can also use enable the floating action button that stays visible as readers scroll. Shortcode support is coming in a future release.

    = Will this plugin slow down my site? =

    No. The plugin uses lightweight, dependency-free JavaScript and CSS. Summary generation only happens when a reader clicks a button — nothing loads on page render that would affect performance.

    = Is the plugin accessible? =

    Yes. Pointwise Summary follows WCAG 2.1 Level AA standards with full keyboard navigation, ARIA labels, semantic HTML, and clear focus indicators.

    = Does it work with my theme? =

    Yes. The plugin uses native WordPress Gutenberg components and follows WordPress coding standards. It is compatible with all properly built themes, including Full Site Editing (FSE) themes.

    = Is it mobile-friendly? =

    Yes. All buttons and layouts are fully responsive and adapt automatically to phone, tablet, and desktop screen sizes.

    = Can I translate the button text? =

    Yes. All button text is fully customizable and supports WordPress translation functions. The plugin is translation-ready for multi-language websites, with summarization support for 50+ languages.

    = Does the plugin include social sharing? =

    Yes. Pointwise Summary has built-in social sharing with support for 7 networks: X / Twitter, Facebook, LinkedIn, Telegram, WhatsApp, Reddit, and Email. Share buttons appear alongside your AI summary buttons in the same row. You can enable or disable each network individually from Display Settings, and X / Twitter supports an optional @mention so shares can tag your account automatically.

    = What is the Floating Action Button? =

    The Floating Action Button (FAB) is a persistent button fixed to a corner of the screen that stays visible as the reader scrolls. It provides access to the summary at any point during reading, without requiring the user to scroll back to the button's original position.

    == Usage ==

    = Basic Setup =

    1. Go to AI Summary (TL;DR) → AI Settings to configure your AI model and API key
    2. Go to AI Summary (TL;DR) → Display Settings to customize button appearance and placement
    3. Save changes and view your posts to see the summary buttons in action

    = Pro Tips =

    For blog posts:

    • Place summary buttons at the top of long articles
    • Use button text like "Don't have time? Get the TL;DR →"
    • Link to a summary section anchored at the bottom of the post

    For documentation:

    • Add "Quick Start" and "Full Guide" buttons side by side
    • Use contrasting colors to make the summary option stand out
    • Try vertical button stacks in sidebars for easier scanning

    For maximum engagement:

    • Use action-oriented labels: "Summarize This", "Get Key Points", "Show TL;DR"
    • Add emoji for visual cues: "⚡ Quick Read" or "📝 Full Article"
    • Test different button positions to find what works best for your audience

    == Changelog ==

    = 1.2.0 2026-06-22 =

    • Added: Custom Logo component for improved branding consistency
    • Enhanced: Admin sidebar UI updated with unified logo integration
    • Added: WordPress.org branding assets (banner and icons)
    • Improved: Plugin description and documentation clarity for better WordPress.org compliance
    • Updated: Plugin tags and metadata for improved discoverability and search relevance
    • Updated: Compatibility flag to WordPress 7.0

    = 1.1.0 =

    • Added: Create Action Button to navigate Setting Page
    • Added: Add screenshot for WordPress readme
    • Added: Create blueprint.json for WordPress.com integration

    = 1.0.0 - 2026-02-10 =

    • Initial stable release with all core features and social sharing support
    • Added: AI settings for six AI models, including ChatGPT, Claude, Gemini, Grok, Perplexity, and Google AI
    • Added: Display settings for automatic insertion, floating action button, collapse Button and visual styles
    • Added: Built-in social sharing with support for 7 networks
    • Added: SEO optimizations with nofollow links and clean link profiles
    • Added: Accessibility features with WCAG 2.1 AA compliance
    • Added: Performance & Caching optimizations with lightweight code and no render-blocking resources
    • Added: Add a custom CSS class to summary buttons for advanced styling
    • Added: Option to choose between <a> links (nofollow) or <button> elements for summary buttons
    • Added: Multilingual Support for 50+ languages with proper internationalization

    = 0.5.0 - 2026-01-20 =

    • Fixed: Improved readme.txt short description for clarity and conciseness
    • Fixed: Streamlined plugin tag list
    • Added: Development section with build instructions

    = 0.4.0 - 2025-12-21 =

    • Enhanced: Block performance and registration system improvements
    • Enhanced: Better compatibility with WordPress 6.8
    • Updated: Documentation and usage examples
    • Fixed: Minor styling inconsistencies across themes
    • Improved: Accessibility features and keyboard navigation

    = 0.3.0 - 2025-02-15 =

    • Added: WordPress 6.7 block manifest system support
    • Enhanced: Faster block loading with optimized registration
    • Improved: Mobile responsiveness and touch interactions
    • Fixed: Button group spacing issues on certain themes

    = 0.2.0 - 2025-01-09 =

    • Updated: Text domain changed to pointwise-summary for proper internationalization
    • Improved: Documentation clarity and TL;DR explanations
    • Enhanced: Plugin directory SEO optimization

    = 0.1.0 - 2024-06-15 =

    • Initial release
    • Added: Horizontal and vertical button layout options
    • Added: Full integration with WordPress core button blocks
    • Added: Comprehensive styling controls (colors, typography, spacing, borders)
    • Added: Block transformations from paragraphs and existing buttons
    • Added: WCAG 2.1 AA accessibility compliance
    • Added: Fully responsive design for all screen sizes
    • Added: Gutenberg Block API v3 support

    == Screenshots ==

    1. Ai-Settings.
    2. Ai-Settings Two.
    3. Display Settings.
    4. Advanced Settings.

    == Support ==

    For help, please use the plugin's support forum on WordPress.org. Include your WordPress version, active theme name, and a description of the issue when posting.

    RoomKit

    Minimum version: 0.7.0 | Python 3.12+

    RoomKit is a pure async Python library for building multi-channel conversation systems. It provides room-based abstractions for managing conversations across SMS, Email, Voice, WebSocket, AI, and other channels with pluggable storage, identity resolution, hooks, and realtime events. Pydantic 2.x, fully typed, zero required dependencies beyond Pydantic.

    RoomKit is a pure async Python library for building multi-channel conversation systems. It provides a room-based abstraction where conversations happen in rooms, participants communicate through channels, and hooks let you intercept and modify the flow.

    Key Concepts

    • Room — A container for a conversation. Holds participants, channel bindings, and an ordered event timeline.
    • Channel — A communication endpoint (SMS, Email, Voice, AI, WebSocket, etc.). Registered once, attached to many rooms.
    • Hook — A function that intercepts events at specific points in the pipeline. Can block, modify, or observe messages.
    • Event Router — Broadcasts events to all channels attached to a room, with content transcoding per channel capabilities.
    • Identity Pipeline — Maps external sender IDs to known participants with challenge/response flows.
    • Realtime Events — Ephemeral events (typing, presence, reactions) that are not stored in history.

    Architecture at a Glance

    RoomKit uses a hub-and-spoke model. The RoomKit orchestrator sits at the center. Channels connect on the edges. Messages flow inbound through a defined pipeline, get stored, then broadcast outbound to all attached channels.

    Inbound Message
      -> InboundRoomRouter.route()       # Find target room
      -> Channel.handle_inbound()        # Parse -> RoomEvent
      -> IdentityResolver.resolve()      # Identify sender
      -> BEFORE_BROADCAST hooks          # Can block/modify
      -> Store event
      -> EventRouter.broadcast()         # Deliver to all channels
        -> Content transcoding           # Adapt per channel capabilities
        -> Rate limiting + retry
      -> AFTER_BROADCAST hooks           # Async side effects
    

    What You Can Build

    • AI-powered support agents across SMS, WhatsApp, and web chat
    • Voice assistants with real-time STT/TTS and interruption handling
    • Multi-agent pipelines where specialized agents hand off conversations
    • Notification systems that bridge channels (SMS + Email + push)
    • Speech-to-speech AI with Gemini Live, OpenAI Realtime, or Grok

    Design Principles

    • Async-first — All I/O is async. No synchronous blocking.
    • Pluggable everything — Storage, identity, routing, AI providers, voice backends — all swappable via ABCs with in-memory defaults.
    • Zero required deps — Only pydantic>=2.9. Everything else is optional extras.
    • Type-safe — Strict mypy, full type hints, Pydantic models throughout.
    • Python 3.12+ — Uses modern syntax (X | None, not Optional[X]).

    Basic Install

    pip install roomkit
    # or with uv
    uv add roomkit
    

    The core library has a single dependency: pydantic>=2.9.

    Optional Extras

    RoomKit uses optional extras for provider-specific dependencies. Install only what you need:

    AI Providers

    pip install roomkit[anthropic]    # Anthropic (Claude)
    pip install roomkit[openai]       # OpenAI (GPT)
    pip install roomkit[gemini]       # Google Gemini
    pip install roomkit[mistral]      # Mistral AI
    pip install roomkit[vllm]         # vLLM local inference (uses openai SDK)
    pip install roomkit[azure]        # Azure OpenAI
    

    Voice — Speech-to-Text

    pip install roomkit[deepgram]     # Deepgram STT (cloud)
    pip install roomkit[sherpa-onnx]  # SherpaOnnx STT (local, offline)
    pip install roomkit[gradium]      # Gradium STT
    pip install roomkit[qwen-asr]     # Qwen3 ASR
    

    Voice — Text-to-Speech

    pip install roomkit[elevenlabs]   # ElevenLabs TTS (cloud)
    pip install roomkit[sherpa-onnx]  # SherpaOnnx TTS (local, offline)
    pip install roomkit[gradium]      # Gradium TTS
    pip install roomkit[qwen-tts]     # Qwen3 TTS
    pip install roomkit[neutts]       # NeuTTS
    

    Voice — Backends

    pip install roomkit[local-audio]  # Local mic/speaker (sounddevice + numpy)
    pip install roomkit[fastrtc]      # FastRTC WebRTC backend
    pip install roomkit[rtp]          # RTP backend
    pip install roomkit[sip]          # SIP backend
    pip install roomkit[webtransport] # WebTransport backend
    

    Voice — Pipeline

    pip install roomkit[webrtc-aec]   # WebRTC echo cancellation
    pip install roomkit[aicoustics]   # ai|coustics denoiser
    pip install roomkit[smart-turn]   # ML-based turn detection
    

    Realtime Voice (Speech-to-Speech)

    pip install roomkit[realtime-openai]   # OpenAI Realtime API
    pip install roomkit[realtime-gemini]   # Google Gemini Live API
    

    Messaging Providers

    pip install roomkit[twilio]            # Twilio SMS/RCS
    pip install roomkit[telegram]          # Telegram Bot API
    pip install roomkit[teams]             # Microsoft Teams (Bot Framework)
    pip install roomkit[whatsapp-personal] # WhatsApp Personal (neonize)
    pip install roomkit[websocket]         # WebSocket source
    pip install roomkit[sse]               # Server-Sent Events source
    

    Storage & Infrastructure

    pip install roomkit[postgres]          # PostgreSQL persistence (asyncpg)
    pip install roomkit[mcp]               # Model Context Protocol tools
    pip install roomkit[opentelemetry]     # OpenTelemetry tracing
    

    Meta Extras

    pip install roomkit[providers]  # All AI + transport providers
    pip install roomkit[sources]    # All event-driven sources
    pip install roomkit[dev]        # Development (test + lint + type check)
    

    Environment Variables

    Provider-specific API keys are passed via configuration objects, not environment variables. Example:

    from roomkit.providers.anthropic.config import AnthropicConfig
    
    config = AnthropicConfig(api_key="sk-ant-...")
    

    For voice providers, use lazy loaders to avoid import-time dependency checks:

    from roomkit.voice import get_deepgram_provider, get_deepgram_config
    
    DeepgramSTTProvider = get_deepgram_provider()
    DeepgramConfig = get_deepgram_config()
    
    stt = DeepgramSTTProvider(DeepgramConfig(api_key="..."))
    

    Development Setup

    git clone https://github.com/roomkit-live/roomkit
    cd roomkit
    uv sync --extra dev    # Install all dev dependencies
    make all               # Run lint + typecheck + security + tests
    

    A complete working example: two WebSocket users chatting with an AI assistant in a moderated room.

    from __future__ import annotations
    
    import asyncio
    
    from roomkit import (
        ChannelCategory,
        HookResult,
        HookTrigger,
        InboundMessage,
        RoomContext,
        RoomEvent,
        RoomKit,
        TextContent,
        WebSocketChannel,
    )
    from roomkit.channels.ai import AIChannel
    from roomkit.providers.ai.mock import MockAIProvider
    
    
    async def main() -> None:
        # 1. Create the framework instance
        kit = RoomKit()
    
        # 2. Create channels
        ws_alice = WebSocketChannel("ws-alice")
        ws_bob = WebSocketChannel("ws-bob")
        ai = AIChannel("ai-assistant", provider=MockAIProvider(responses=["Got it!"]))
    
        # 3. Register channels with the framework
        kit.register_channel(ws_alice)
        kit.register_channel(ws_bob)
        kit.register_channel(ai)
    
        # 4. Wire up receive callbacks (in production, WebSocket sends to clients)
        alice_inbox: list[RoomEvent] = []
        bob_inbox: list[RoomEvent] = []
    
        async def alice_recv(_conn: str, event: RoomEvent) -> None:
            alice_inbox.append(event)
    
        async def bob_recv(_conn: str, event: RoomEvent) -> None:
            bob_inbox.append(event)
    
        ws_alice.register_connection("alice-conn", alice_recv)
        ws_bob.register_connection("bob-conn", bob_recv)
    
        # 5. Create a room and attach channels
        await kit.create_room(room_id="demo-room")
        await kit.attach_channel("demo-room", "ws-alice")
        await kit.attach_channel("demo-room", "ws-bob")
        await kit.attach_channel(
            "demo-room", "ai-assistant", category=ChannelCategory.INTELLIGENCE
        )
    
        # 6. Add a BEFORE_BROADCAST hook for content moderation
        @kit.hook(HookTrigger.BEFORE_BROADCAST, name="profanity_filter")
        async def profanity_filter(event: RoomEvent, ctx: RoomContext) -> HookResult:
            if isinstance(event.content, TextContent) and "badword" in event.content.body:
                return HookResult.block("Message blocked by profanity filter")
            return HookResult.allow()
    
        # 7. Send messages through the inbound pipeline
        result = await kit.process_inbound(
            InboundMessage(
                channel_id="ws-alice",
                sender_id="alice",
                content=TextContent(body="Hello everyone!"),
            )
        )
        print(f"Alice sent 'Hello everyone!' -> blocked={result.blocked}")
    
        # Both Bob and the AI receive Alice's message.
        # The AI responds with "Got it!" which is broadcast to Alice and Bob.
    
        # 8. Query stored conversation history
        events = await kit.store.list_events("demo-room")
        for ev in events:
            print(f"  [{ev.source.channel_id}] {ev.content.body}")
    
    
    if __name__ == "__main__":
        asyncio.run(main())
    

    What Just Happened

    1. RoomKit() created the framework with in-memory defaults (store, locks, realtime).
    2. Channels were registered globally, then attached to a room.
    3. AI channel was attached with category=ChannelCategory.INTELLIGENCE — it receives messages and generates responses.
    4. A BEFORE_BROADCAST hook runs synchronously before every broadcast. It can block(), allow(), or modify() events.
    5. process_inbound() ran the full pipeline: route -> parse -> identity -> hooks -> store -> broadcast.
    6. The AI's response was automatically routed back through the same pipeline (with chain depth tracking to prevent loops).

    Core Pattern

    Every RoomKit application follows this pattern:

    from roomkit import RoomKit
    
    # 1. Create framework
    kit = RoomKit()
    
    # 2. Register channels
    kit.register_channel(channel)
    
    # 3. Create rooms and attach channels
    await kit.create_room(room_id="my-room")
    await kit.attach_channel("my-room", "channel-id")
    
    # 4. Process inbound messages
    await kit.process_inbound(InboundMessage(...))
    

    Next Steps

    • Add hooks for moderation, logging, or analytics (see Hooks)
    • Configure AI providers for real LLM responses (see AI Channels)
    • Add voice with STT/TTS (see Voice Channels)
    • Set up multi-agent orchestration (see Orchestration)
    • Deploy with PostgreSQL storage (see Storage)

    Hub-and-Spoke Model

    RoomKit uses a hub-and-spoke architecture. The RoomKit class is the central hub. Channels are the spokes. Messages flow in through channels, get processed by the hub, then broadcast out to all channels attached to the room.

                        +------------------+
                        |     RoomKit      |
                        |  (orchestrator)  |
                        +--------+---------+
                                 |
             +-------------------+-------------------+
             |         |         |         |         |
         +---+---+ +---+---+ +--+---+ +---+---+ +---+---+
         |  SMS  | | Email | | Voice| |  AI   | |  WS   |
         +-------+ +-------+ +------+ +-------+ +-------+
    

    Inbound Pipeline

    Every inbound message follows an immutable pipeline order:

    1. InboundRoomRouter.route()        # Find target room by channel binding
    2. Channel.handle_inbound()         # Parse external format -> RoomEvent
    3. IdentityResolver.resolve()       # Map sender_id -> participant
    4. Identity hooks                   # ON_IDENTITY_AMBIGUOUS / ON_IDENTITY_UNKNOWN
    5. Room lock acquired               # Per-room atomic processing
    6. Idempotency check                # Deduplicate by provider_message_id
    7. BEFORE_BROADCAST hooks           # Sync: can block or modify the event
    8. Store event + update counters    # Persist to ConversationStore
    9. EventRouter.broadcast()          # Deliver to all attached channels
       -> Content transcoding           # Adapt content per channel capabilities
       -> Rate limiting                 # TokenBucketRateLimiter per binding
       -> Retry with backoff            # RetryPolicy per binding
    10. AFTER_BROADCAST hooks           # Async: fire-and-forget side effects
    11. Room lock released
    

    This order is defined in the RFC (Section 10.1) and must not be reordered.

    Channel Categories

    Channels have two categories:

    • TRANSPORT — Push messages to users (SMS, Email, WebSocket, Voice). Default category.
    • INTELLIGENCE — Generate responses (AI, agents). Receives broadcasts, responds through the inbound pipeline.

    Channel-to-AI Message Flow (Reentry Loop)

    When you send a message from a transport channel, it flows to the AI and back automatically. Here's the complete flow:

    1. Transport channel (SMS/WS/Voice) → kit.process_inbound(InboundMessage)
    2. Inbound pipeline: route → parse → identity → hooks → store
    3. EventRouter.broadcast() → delivers event to ALL attached channels
       ├── Transport channels: note delivery (no response)
       └── AI channel: calls LLM provider → generates response
           └── Returns ChannelOutput(response_events=[RoomEvent])
    4. REENTRY LOOP: AI response re-enters as new inbound event
       ├── BEFORE_BROADCAST hooks run again (ConversationRouter stamps routing)
       ├── Event stored in timeline
       ├── Broadcast to all channels again
       │   ├── Transport channels: DELIVER the AI response to users
       │   └── Other AI channels: see the response (may generate follow-up)
       └── If follow-up AI responses exist → loop again (chain depth checked)
    5. Chain depth limit (default max=5) → stops AI-to-AI infinite loops
    6. AFTER_BROADCAST hooks fire (async side effects)
    

    Key insight: The "pipeline" is a reentry loop — not a linear chain. AI responses go back through the same BEFORE_BROADCAST hooks, content transcoding, and broadcast cycle as user messages.

    Connecting Channel to AI — Minimal Example

    from roomkit import RoomKit, WebSocketChannel, ChannelCategory, InboundMessage, TextContent
    from roomkit.channels.ai import AIChannel
    from roomkit.providers.anthropic.ai import AnthropicAIProvider
    from roomkit.providers.anthropic.config import AnthropicConfig
    
    kit = RoomKit()
    
    # Transport channel (user-facing)
    ws = WebSocketChannel("ws-user")
    ws.register_connection("conn-1", on_recv)
    kit.register_channel(ws)
    
    # Intelligence channel (AI)
    ai = AIChannel("ai", provider=AnthropicAIProvider(AnthropicConfig(
        api_key="sk-ant-...", model="claude-sonnet-4-20250514",
    )))
    kit.register_channel(ai)
    
    # Room wires them together
    await kit.create_room(room_id="chat")
    await kit.attach_channel("chat", "ws-user")  # TRANSPORT (default)
    await kit.attach_channel("chat", "ai", category=ChannelCategory.INTELLIGENCE)
    
    # User sends message → AI responds → response delivered to WebSocket
    await kit.process_inbound(
        InboundMessage(channel_id="ws-user", sender_id="user", content=TextContent(body="Hi!"))
    )
    # on_recv callback fires with the AI's response
    

    Pipeline vs Pipeline

    RoomKit uses "Pipeline" in two contexts — don't confuse them:

    TermWhat It IsWhere
    Inbound processing pipelineThe 11-step message processing flow (route → parse → identity → hooks → store → broadcast)core/mixins/inbound_locked.py
    Pipeline orchestration strategyA linear agent chain (triage → handler → resolver) for multi-agent handoffsfrom roomkit import Pipeline

    The inbound pipeline processes every message. The Pipeline strategy controls which agent handles each turn.

    Pluggable Components

    Every core component follows the ABC + default pattern:

    ComponentABCDefaultPurpose
    ConversationStorestore/base.pyInMemoryStoreRoom, event, participant persistence
    RoomLockManagercore/locks.pyInMemoryLockManagerPer-room atomic processing
    RealtimeBackendrealtime/base.pyInMemoryRealtimeEphemeral events (typing, presence)
    IdentityResolveridentity/base.pyNone (disabled)Sender -> participant mapping
    InboundRoomRoutercore/inbound_router.pyDefaultInboundRoomRouterRoute messages to rooms

    Replace any component at construction:

    from roomkit import RoomKit
    from roomkit.store.postgres import PostgresStore
    
    kit = RoomKit(store=PostgresStore("postgresql://..."))
    

    Room Lifecycle

    ACTIVE -> PAUSED -> CLOSED -> ARCHIVED
               ^          |
               +----------+  (can close from paused)
    
    • ACTIVE — Accepting messages, all channels active.
    • PAUSED — Messages queued, channels paused. Auto-transition via inactive_after_seconds.
    • CLOSED — No new messages. Auto-transition via closed_after_seconds.
    • ARCHIVED — Terminal state, read-only.

    Event Model

    Every message becomes a RoomEvent with:

    • content — One of 11 content types (TextContent, RichContent, MediaContent, AudioContent, VideoContent, LocationContent, CompositeContent, TemplateContent, EditContent, DeleteContent, SystemContent)
    • source — Who sent it (channel_id, participant_id, direction)
    • index — Sequential, monotonically increasing per room
    • metadata — Arbitrary key-value data

    Voice Architecture

    The voice subsystem has three layers:

    1. VoiceBackend — Pure audio transport (mic, SIP, RTP, WebRTC). No speech detection.
    2. AudioPipeline — Processes audio frames: Resampler -> Recorder -> AEC -> AGC -> Denoiser -> VAD -> Diarization.
    3. VoiceChannel — Wires backend -> pipeline -> STT/TTS, handles interruption and turn detection.
    Inbound:   Backend -> [Resampler] -> [Recorder] -> [AEC] -> [AGC] -> [Denoiser] -> VAD -> [Diarization] + [DTMF]
    Outbound:  TTS -> [PostProcessors] -> [Recorder] -> AEC.feed_reference -> [Resampler] -> Backend
    

    Realtime Voice Architecture

    Speech-to-speech AI (Gemini Live, OpenAI Realtime) bypasses STT/TTS entirely:

    • RealtimeVoiceProvider — Handles the AI model connection (WebSocket to Gemini/OpenAI).
    • RealtimeAudioTransport — Handles browser/client audio (WebSocket or WebRTC).
    • RealtimeVoiceChannel — Bridges transport <-> provider with tool calling support.

    Creating Rooms

    from roomkit import RoomKit
    
    kit = RoomKit()
    
    # Auto-generated ID
    room = await kit.create_room()
    
    # Explicit ID with metadata
    room = await kit.create_room(
        room_id="support-123",
        metadata={"topic": "billing", "priority": "high"},
    )
    

    Room Lifecycle

    from roomkit import RoomKit
    
    kit = RoomKit()
    
    # Create room with auto-pause and auto-close timers
    room = await kit.create_room(room_id="session-1")
    
    # Update metadata
    await kit.update_room_metadata("session-1", {"status": "escalated"})
    
    # Close a room
    await kit.close_room("session-1")
    
    # Check timers (call periodically, e.g. every 60s)
    transitioned = await kit.check_all_timers()
    

    Channel Types

    RoomKit ships with these channel types:

    ChannelFactory / ClassCategoryUse Case
    SMSSMSChannel(id, provider=...)TRANSPORTText messages via Twilio, Telnyx, Sinch
    RCSRCSChannel(id, provider=...)TRANSPORTRich messaging via Twilio, Telnyx
    EmailEmailChannel(id, provider=...)TRANSPORTEmail via ElasticEmail, SendGrid
    WhatsAppWhatsAppChannel(id, provider=...)TRANSPORTWhatsApp Business API
    WhatsApp PersonalWhatsAppPersonalChannel(id, provider=...)TRANSPORTWhatsApp via neonize
    MessengerMessengerChannel(id, provider=...)TRANSPORTFacebook Messenger
    TelegramTelegramChannel(id, provider=...)TRANSPORTTelegram Bot API
    TeamsTeamsChannel(id, provider=...)TRANSPORTMicrosoft Teams Bot Framework
    HTTPHTTPChannel(id, provider=...)TRANSPORTGeneric webhook
    WebSocketWebSocketChannel(id)TRANSPORTReal-time bidirectional
    AIAIChannel(id, provider=...)INTELLIGENCELLM responses
    AgentAgent(id, provider=...)INTELLIGENCEAgent with tools + greeting
    VoiceVoiceChannel(id, stt=..., tts=..., backend=...)TRANSPORTReal-time audio
    Realtime VoiceRealtimeVoiceChannel(id, provider=...)INTELLIGENCESpeech-to-speech AI

    Registering and Attaching Channels

    Channels are registered globally, then attached to specific rooms:

    from roomkit import RoomKit, SMSChannel, ChannelCategory
    from roomkit.channels.ai import AIChannel
    from roomkit.providers.twilio.sms import TwilioSMSProvider
    from roomkit.providers.twilio.config import TwilioConfig
    from roomkit.providers.anthropic.ai import AnthropicAIProvider
    from roomkit.providers.anthropic.config import AnthropicConfig
    
    kit = RoomKit()
    
    # Create and register channels
    sms = SMSChannel("sms-main", provider=TwilioSMSProvider(TwilioConfig(
        account_sid="AC...",
        auth_token="...",
        from_number="+1234567890",
    )))
    ai = AIChannel("ai-agent", provider=AnthropicAIProvider(AnthropicConfig(
        api_key="sk-ant-...",
        model="claude-sonnet-4-20250514",
    )))
    
    kit.register_channel(sms)
    kit.register_channel(ai)
    
    # Create room and attach
    await kit.create_room(room_id="support-room")
    await kit.attach_channel("support-room", "sms-main")
    await kit.attach_channel("support-room", "ai-agent", category=ChannelCategory.INTELLIGENCE)
    

    Channel Access Levels

    Control what a channel can do in a room:

    from roomkit import Access
    
    # Default: read and write
    await kit.attach_channel("room", "channel", access=Access.READ_WRITE)
    
    # Read only: receives messages but cannot send
    await kit.attach_channel("room", "channel", access=Access.READ_ONLY)
    
    # Write only: sends messages but doesn't receive broadcasts
    await kit.attach_channel("room", "channel", access=Access.WRITE_ONLY)
    
    # None: temporarily disabled
    await kit.set_access("room", "channel", Access.NONE)
    

    Muting Channels

    Muting suppresses a channel's output (AI responses) without detaching it. Side effects (tasks, observations) still fire.

    # Mute AI output
    await kit.mute("room", "ai-agent")
    
    # Unmute
    await kit.unmute("room", "ai-agent")
    
    # Mute only the output direction (AI won't respond, but still sees messages)
    await kit.mute_output("room", "ai-agent")
    await kit.unmute_output("room", "ai-agent")
    

    Per-Room Channel Configuration

    Pass metadata when attaching to customize per-room behavior:

    await kit.attach_channel(
        "weather-room",
        "ai-agent",
        category=ChannelCategory.INTELLIGENCE,
        metadata={
            "system_prompt": "You are a weather assistant.",
            "temperature": 0.3,
            "tools": [
                {
                    "name": "get_weather",
                    "description": "Get current weather for a city",
                    "parameters": {
                        "type": "object",
                        "properties": {
                            "city": {"type": "string"},
                        },
                        "required": ["city"],
                    },
                },
            ],
        },
    )
    

    Content Types

    Events carry typed content. RoomKit supports 11 content types, all discriminated unions on the type field:

    Typetype LiteralKey FieldsUse Case
    TextContent"text"body, languagePlain text messages
    RichContent"rich"body, format (html/markdown), plain_text, buttons, cards, quick_repliesFormatted messages with UI elements
    MediaContent"media"url, mime_type, filename, size_bytes, captionImages, documents, files
    AudioContent"audio"url, mime_type, duration_seconds, transcriptVoice messages
    VideoContent"video"url, mime_type, duration_seconds, thumbnail_urlVideo messages
    LocationContent"location"latitude, longitude, label, addressGeographic coordinates
    CompositeContent"composite"parts (list of EventContent, max depth 5)Multi-part messages
    TemplateContent"template"template_id, language, parameters, bodyWhatsApp Business / RCS templates
    SystemContent"system"body, code, data (dict)System-generated events
    EditContent"edit"target_event_id, new_content, edit_sourceMessage edits
    DeleteContent"delete"target_event_id, delete_type, reasonMessage deletion
    from roomkit.models.event import (
        TextContent,
        RichContent,
        MediaContent,
        AudioContent,
        VideoContent,
        LocationContent,
        CompositeContent,
        TemplateContent,
        SystemContent,
        EditContent,
        DeleteContent,
    )
    
    # Plain text
    TextContent(body="Hello!")
    
    # Rich content with buttons and Markdown
    RichContent(
        body="Choose an option:",
        format="markdown",
        plain_text="Choose an option:",
        buttons=[{"text": "Option A", "payload": "a"}, {"text": "Option B", "payload": "b"}],
    )
    
    # Media attachment
    MediaContent(url="https://example.com/image.png", mime_type="image/png", caption="Product photo")
    
    # Audio with transcript
    AudioContent(url="https://example.com/voice.ogg", mime_type="audio/ogg", transcript="Hello there")
    
    # Location
    LocationContent(latitude=40.7128, longitude=-74.0060, label="NYC Office", address="New York, NY")
    
    # Template (WhatsApp Business)
    TemplateContent(template_id="order_confirmation", language="en", parameters={"1": "ORD-123"})
    
    # Multi-part message (max depth 5)
    CompositeContent(parts=[
        TextContent(body="Here's the photo:"),
        MediaContent(url="https://example.com/photo.jpg", mime_type="image/jpeg"),
    ])
    
    # Edit a previous message
    EditContent(target_event_id="evt-abc", new_content=TextContent(body="Corrected text"))
    
    # Delete a message
    DeleteContent(target_event_id="evt-abc", delete_type="sender", reason="User retracted")
    

    Content Transcoding

    When broadcasting events, the EventRouter automatically transcodes content to match each target channel's capabilities. A channel that only supports TEXT will receive a text fallback of a RichContent message.

    How Transcoding Works

    1. Event broadcast starts → EventRouter iterates over all target channel bindings
    2. For each target, call ContentTranscoder.transcode(content, source_binding, target_binding)
    3. Transcoder checks target_binding.capabilities.media_types to decide if content passes through or needs conversion
    4. If transcode returns None → delivery is skipped for that channel (content cannot be represented)

    Channel Capabilities

    Each channel binding declares what content types it supports via ChannelCapabilities:

    from roomkit.models.channel import ChannelCapabilities
    from roomkit.models.enums import ChannelMediaType
    
    # SMS channel: text only, 160 char limit
    sms_caps = ChannelCapabilities(
        media_types=[ChannelMediaType.TEXT],
        max_length=160,
    )
    
    # WebSocket channel: rich content, media, audio, video
    ws_caps = ChannelCapabilities(
        media_types=[
            ChannelMediaType.TEXT,
            ChannelMediaType.RICH,
            ChannelMediaType.MEDIA,
            ChannelMediaType.AUDIO,
            ChannelMediaType.VIDEO,
            ChannelMediaType.LOCATION,
        ],
        supports_edit=True,
        supports_delete=True,
        supports_reactions=True,
        supports_typing=True,
    )
    

    ChannelMediaType enum values: TEXT, RICH, MEDIA, AUDIO, VIDEO, LOCATION, TEMPLATE.

    Default Fallback Chain

    The DefaultContentTranscoder applies these fallback rules:

    Content TypeIf Target Supports ItFallback
    TextContentAlways passes through—
    RichContentRICH in media_types → passTextContent(plain_text or body)
    MediaContentMEDIA in media_types → passTextContent("[Media: {caption or filename or url}]")
    AudioContentAUDIO in media_types → passTextContent(transcript) or "[Voice message: {url}]"
    VideoContentVIDEO in media_types → passTextContent("[Video: {url}]")
    LocationContentLOCATION in media_types → passTextContent("[Location: {label} ({lat}, {lon})]")
    CompositeContentRecursive transcode of all partsIf all parts become text → flatten to single TextContent
    TemplateContentTEMPLATE in media_types → passTextContent(body or "[Template: {id}]")
    EditContentsupports_edit → passTranscode new_content + prefix "Correction:"
    DeleteContentsupports_delete → passTextContent("[Message deleted]")

    Custom Content Transcoder

    Implement the ContentTranscoder ABC to customize how content is adapted:

    from roomkit.core.router import ContentTranscoder
    from roomkit.models.channel import ChannelBinding
    from roomkit.models.event import EventContent, TextContent, MediaContent
    from roomkit.models.enums import ChannelMediaType
    
    class MyTranscoder(ContentTranscoder):
        async def transcode(
            self,
            content: EventContent,
            source_binding: ChannelBinding,
            target_binding: ChannelBinding,
        ) -> EventContent | None:
            target_types = target_binding.capabilities.media_types
    
            # Custom: convert images to descriptive text for voice channels
            if isinstance(content, MediaContent) and content.mime_type.startswith("image/"):
                if ChannelMediaType.MEDIA not in target_types:
                    caption = content.caption or "an image"
                    return TextContent(body=f"[Image received: {caption}]")
    
            # Custom: truncate long text for SMS
            if isinstance(content, TextContent):
                max_len = target_binding.capabilities.max_length
                if max_len and len(content.body) > max_len:
                    return TextContent(body=content.body[: max_len - 3] + "...")
    
            # Fall through to default behavior for other types
            return content
    
    # Override the default transcoder on the RoomKit instance
    kit = RoomKit()
    kit._transcoder = MyTranscoder()
    

    Multichannel Transcoding Example

    A single event gets adapted differently for each channel:

    # User sends a rich message with an image from WebSocket
    content = CompositeContent(parts=[
        RichContent(body="**Check this out!**", format="markdown", plain_text="Check this out!"),
        MediaContent(url="https://example.com/photo.jpg", mime_type="image/jpeg", caption="Sunset"),
    ])
    
    # WebSocket channel: receives CompositeContent as-is (supports RICH + MEDIA)
    # SMS channel: receives TextContent("Check this out!\n[Media: Sunset]")
    # Voice channel (TTS): receives TextContent("Check this out!\n[Media: Sunset]")
    

    Detaching Channels

    await kit.detach_channel("room", "channel-id")
    

    Binding Metadata Updates

    await kit.update_binding_metadata("room", "ai-agent", {"temperature": 0.9})
    

    Querying Rooms

    # Get a room
    room = await kit.get_room("room-id")
    
    # List bindings
    bindings = await kit.store.list_bindings("room-id")
    
    # List participants
    participants = await kit.store.list_participants("room-id")
    
    # Query event timeline
    events = await kit.get_timeline("room-id", offset=0, limit=50)
    

    Hooks intercept events at specific points in the processing pipeline. They can block messages, modify content, trigger side effects, or observe the conversation.

    Hook Basics

    from roomkit import RoomKit, HookTrigger, HookExecution, HookResult, RoomEvent, RoomContext
    
    kit = RoomKit()
    
    # Sync hook: runs BEFORE broadcast, can block or modify
    @kit.hook(HookTrigger.BEFORE_BROADCAST)
    async def content_filter(event: RoomEvent, ctx: RoomContext) -> HookResult:
        if "spam" in event.content.body.lower():
            return HookResult.block("Spam detected")
        return HookResult.allow()
    
    # Async hook: runs AFTER broadcast, fire-and-forget
    @kit.hook(HookTrigger.AFTER_BROADCAST, execution=HookExecution.ASYNC)
    async def log_event(event: RoomEvent, ctx: RoomContext) -> None:
        await analytics.track("message", {"room": event.room_id})
    

    HookResult

    Sync hooks (BEFORE_BROADCAST) must return a HookResult:

    from roomkit import HookResult, TextContent
    
    # Allow the event to proceed
    HookResult.allow()
    
    # Block the event with a reason
    HookResult.block("Contains prohibited content")
    
    # Modify the event before broadcast
    modified = event.model_copy(update={"content": TextContent(body="[REDACTED]")})
    HookResult.modify(modified)
    

    Hook Priority

    Lower priority numbers run first. Default is 0.

    # Runs first (priority=0)
    @kit.hook(HookTrigger.BEFORE_BROADCAST, name="profanity_filter", priority=0)
    async def profanity_filter(event: RoomEvent, ctx: RoomContext) -> HookResult:
        blocked_words = {"badword", "spam", "scam"}
        if isinstance(event.content, TextContent):
            words = set(event.content.body.lower().split())
            if words & blocked_words:
                return HookResult.block(f"Blocked: {words & blocked_words}")
        return HookResult.allow()
    
    # Runs second (priority=1)
    @kit.hook(HookTrigger.BEFORE_BROADCAST, name="pii_redactor", priority=1)
    async def pii_redactor(event: RoomEvent, ctx: RoomContext) -> HookResult:
        import re
        if isinstance(event.content, TextContent):
            redacted = re.sub(
                r"\b\d{3}[-.]?\d{3}[-.]?\d{4}\b", "[REDACTED]", event.content.body
            )
            if redacted != event.content.body:
                modified = event.model_copy(update={"content": TextContent(body=redacted)})
                return HookResult.modify(modified)
        return HookResult.allow()
    

    Hook Filters

    Filter hooks by channel type, channel ID, or direction:

    from roomkit.models.enums import ChannelType, ChannelDirection
    
    @kit.hook(
        HookTrigger.AFTER_BROADCAST,
        execution=HookExecution.ASYNC,
        channel_types={ChannelType.SMS},
        directions={ChannelDirection.INBOUND},
        priority=10,
    )
    async def sms_audit(event: RoomEvent, ctx: RoomContext) -> None:
        await audit_log.record(event)
    

    Room-Scoped Hooks

    Add hooks to specific rooms instead of globally:

    from roomkit import HookExecution
    
    await kit.add_room_hook(
        room_id="vip-room",
        trigger=HookTrigger.BEFORE_BROADCAST,
        execution=HookExecution.SYNC,
        fn=my_hook_function,
        name="vip_filter",
    )
    
    # Remove later
    await kit.remove_room_hook("vip-room", "vip_filter")
    

    Complete Hook Trigger Reference

    Event Pipeline

    TriggerExecutionSignatureDescription
    BEFORE_BROADCASTSYNC(event, ctx) -> HookResultBefore event is stored and broadcast. Can block/modify.
    AFTER_BROADCASTASYNC(event, ctx) -> NoneAfter event is broadcast. Fire-and-forget side effects.

    Channel Lifecycle

    TriggerExecutionSignatureDescription
    ON_CHANNEL_ATTACHEDASYNC(event, ctx) -> NoneChannel was attached to a room
    ON_CHANNEL_DETACHEDASYNC(event, ctx) -> NoneChannel was detached from a room
    ON_CHANNEL_MUTEDASYNC(event, ctx) -> NoneChannel was muted in a room
    ON_CHANNEL_UNMUTEDASYNC(event, ctx) -> NoneChannel was unmuted in a room

    Room Lifecycle

    TriggerExecutionSignatureDescription
    ON_ROOM_CREATEDASYNC(event, ctx) -> NoneRoom was created
    ON_ROOM_PAUSEDASYNC(event, ctx) -> NoneRoom was paused
    ON_ROOM_CLOSEDASYNC(event, ctx) -> NoneRoom was closed

    Identity

    TriggerExecutionSignatureDescription
    ON_IDENTITY_AMBIGUOUSSYNC(event, ctx) -> IdentityHookResultMultiple identity matches found
    ON_IDENTITY_UNKNOWNSYNC(event, ctx) -> IdentityHookResultNo identity match found
    ON_PARTICIPANT_IDENTIFIEDASYNC(event, ctx) -> NoneParticipant was successfully identified

    Delivery

    TriggerExecutionSignatureDescription
    ON_DELIVERY_STATUSASYNC(status) -> NoneDelivery status update from provider

    Side Effects

    TriggerExecutionSignatureDescription
    ON_TASK_CREATEDASYNC(event, ctx) -> NoneAI extracted a task from conversation
    ON_ERRORASYNC(event, ctx) -> NoneError occurred during processing

    Voice

    TriggerExecutionSignatureDescription
    ON_SPEECH_STARTASYNC(event, ctx) -> NoneVAD detected speech start
    ON_SPEECH_ENDASYNC(event, ctx) -> NoneVAD detected speech end
    ON_TRANSCRIPTIONASYNC(event, ctx) -> NoneSTT produced a transcription
    BEFORE_TTSSYNC(event, ctx) -> HookResultBefore text is sent to TTS. Can block/modify.
    AFTER_TTSASYNC(event, ctx) -> NoneAfter TTS audio is generated

    Voice Pipeline

    TriggerExecutionSignatureDescription
    ON_VAD_SILENCEASYNC(event, ctx) -> NoneVAD detected silence
    ON_VAD_AUDIO_LEVELASYNC(event, ctx) -> NoneAudio level update from VAD
    ON_SPEAKER_CHANGEASYNC(event, ctx) -> NoneDiarization detected speaker change
    ON_BARGE_INASYNC(event, ctx) -> NoneUser interrupted TTS playback
    ON_TTS_CANCELLEDASYNC(event, ctx) -> NoneTTS playback was cancelled
    ON_DTMFASYNC(event, ctx) -> NoneDTMF tone detected
    ON_TURN_COMPLETEASYNC(event, ctx) -> NoneTurn detector says turn is complete
    ON_TURN_INCOMPLETEASYNC(event, ctx) -> NoneTurn detector says turn is incomplete
    ON_BACKCHANNELASYNC(event, ctx) -> NoneBackchannel detected (uh-huh, yeah)
    ON_RECORDING_STARTEDASYNC(event, ctx) -> NoneAudio recording started
    ON_RECORDING_STOPPEDASYNC(event, ctx) -> NoneAudio recording stopped
    ON_INPUT_AUDIO_LEVELASYNC(event, ctx) -> NoneInput audio level update
    ON_OUTPUT_AUDIO_LEVELASYNC(event, ctx) -> NoneOutput audio level update

    Tool Execution

    TriggerExecutionSignatureDescription
    ON_TOOL_CALLASYNC(event, ctx) -> NoneAI invoked a tool (fires from AIChannel and RealtimeVoiceChannel)

    Realtime Voice

    TriggerExecutionSignatureDescription
    ON_REALTIME_TEXT_INJECTEDASYNC(event, ctx) -> NoneText was injected into realtime session

    Observability

    TriggerExecutionSignatureDescription
    ON_FEEDBACKASYNC(Observation, ctx) -> NoneUser submitted quality feedback via kit.submit_feedback()

    Framework Events

    Framework events are lightweight lifecycle notifications (not message events):

    @kit.on("room_created")
    async def on_room_created(event):
        print(f"Room created: {event.data['room_id']}")
    
    @kit.on("voice_session_started")
    async def on_voice(event):
        print(f"Voice session: {event.data['session_id']}")
    

    Available framework event types: room_created, room_closed, room_paused, room_channel_attached, room_channel_detached, channel_connected, channel_disconnected, voice_session_started, voice_session_ended, source_attached, source_detached, source_error, source_exhausted.

    AIChannel connects rooms to LLM providers. When a message is broadcast to an AI channel, it generates a response using conversation history and re-enters it through the inbound pipeline.

    Basic Setup

    from roomkit import RoomKit, ChannelCategory
    from roomkit.channels.ai import AIChannel
    from roomkit.providers.anthropic.ai import AnthropicAIProvider
    from roomkit.providers.anthropic.config import AnthropicConfig
    
    kit = RoomKit()
    
    ai = AIChannel(
        "ai-assistant",
        provider=AnthropicAIProvider(AnthropicConfig(
            api_key="sk-ant-...",
            model="claude-sonnet-4-20250514",
        )),
        system_prompt="You are a helpful customer support agent.",
        temperature=0.7,
    )
    kit.register_channel(ai)
    
    await kit.create_room(room_id="support")
    await kit.attach_channel("support", "ai-assistant", category=ChannelCategory.INTELLIGENCE)
    

    AI Providers

    ProviderClassConfigExtra
    Anthropic (Claude)AnthropicAIProviderAnthropicConfigroomkit[anthropic]
    OpenAI (GPT)OpenAIAIProviderOpenAIConfigroomkit[openai]
    Google GeminiGeminiAIProviderGeminiConfigroomkit[gemini]
    MistralMistralAIProviderMistralConfigroomkit[mistral]
    Azure OpenAIAzureAIProviderAzureAIConfigroomkit[azure]
    vLLM (local)create_vllm_provider()VLLMConfigroomkit[vllm]
    Mock (testing)MockAIProvider—built-in
    # OpenAI
    from roomkit.providers.openai.ai import OpenAIAIProvider
    from roomkit.providers.openai.config import OpenAIConfig
    
    provider = OpenAIAIProvider(OpenAIConfig(api_key="sk-...", model="gpt-4o"))
    
    # Gemini
    from roomkit.providers.gemini.ai import GeminiAIProvider
    from roomkit.providers.gemini.config import GeminiConfig
    
    provider = GeminiAIProvider(GeminiConfig(api_key="...", model="gemini-2.0-flash"))
    
    # Mock (for testing)
    from roomkit.providers.ai.mock import MockAIProvider
    
    provider = MockAIProvider(responses=["Hello!", "How can I help?"])
    

    Agent Class

    Agent extends AIChannel with role, description, greeting, and memory support — designed for multi-agent orchestration:

    from roomkit import Agent
    from roomkit.providers.ai.mock import MockAIProvider
    
    agent = Agent(
        "support-agent",
        provider=MockAIProvider(responses=["I can help with that."]),
        role="Customer support specialist",
        description="Handles billing and account questions",
        system_prompt="You are a support specialist. Be concise and helpful.",
        greeting="Hi! How can I help you today?",
    )
    

    Tool Calling

    Define tools as JSON schema and attach them to the AI channel:

    from roomkit import ChannelCategory
    from roomkit.channels.ai import AIChannel
    from roomkit.providers.openai.ai import OpenAIAIProvider
    from roomkit.providers.openai.config import OpenAIConfig
    
    ai = AIChannel(
        "ai-assistant",
        provider=OpenAIAIProvider(OpenAIConfig(api_key="sk-...", model="gpt-4o")),
        system_prompt="You help users check the weather.",
        tools=[
            {
                "name": "get_weather",
                "description": "Get current weather for a city",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "city": {"type": "string", "description": "City name"},
                        "units": {"type": "string", "enum": ["celsius", "fahrenheit"]},
                    },
                    "required": ["city"],
                },
            },
        ],
    )
    

    Tool Handler

    Register a handler to execute tool calls via the constructor:

    async def handle_tools(name: str, arguments: dict) -> str:
        if name == "get_weather":
            city = arguments["city"]
            return f'{{"temperature": 22, "condition": "sunny", "city": "{city}"}}'
        return '{"error": "Unknown tool"}'
    
    ai = AIChannel(
        "ai-assistant",
        provider=provider,
        tools=[...],
        tool_handler=handle_tools,
    )
    

    Tool Protocol (Tool ABC)

    For structured tool definitions, use the Tool base class:

    from roomkit.tools.base import Tool
    
    class GetWeather(Tool):
        name = "get_weather"
        description = "Get current weather for a city"
        parameters = {
            "type": "object",
            "properties": {
                "city": {"type": "string"},
            },
            "required": ["city"],
        }
    
        async def execute(self, arguments: dict) -> str:
            return '{"temperature": 22, "condition": "sunny"}'
    
    ai = AIChannel("ai", provider=provider, tools=[GetWeather()])
    

    MCP Tool Provider

    Integrate Model Context Protocol servers:

    from roomkit.tools.mcp import MCPToolProvider
    
    mcp = MCPToolProvider(server_command=["uvx", "mcp-server-sqlite", "--db", "data.db"])
    await mcp.initialize()
    
    ai = AIChannel("ai", provider=provider, tools=mcp.tools())
    

    Per-Room Configuration

    Override AI settings per room via binding metadata:

    await kit.attach_channel(
        "billing-room",
        "ai-agent",
        category=ChannelCategory.INTELLIGENCE,
        metadata={
            "system_prompt": "You are a billing specialist.",
            "temperature": 0.3,
            "tools": [...],
        },
    )
    

    Streaming

    AIChannel supports streaming responses to WebSocket clients:

    from roomkit import WebSocketChannel
    
    ws = WebSocketChannel("ws-user")
    
    # Register with stream support
    ws.register_connection("conn-1", on_recv, stream_send_fn=on_stream)
    
    async def on_stream(conn_id: str, msg) -> None:
        # StreamStart, StreamChunk, StreamEnd
        print(f"Stream: {msg}")
    

    AI Thinking/Reasoning

    Some providers support extended thinking:

    ai = AIChannel(
        "ai",
        provider=AnthropicAIProvider(AnthropicConfig(
            api_key="...",
            model="claude-sonnet-4-20250514",
        )),
        system_prompt="Think step by step.",
        thinking_budget=4096,  # Setting a budget enables thinking mode
    )
    

    Vision Support

    AI providers that support vision can process images sent as MediaContent:

    from roomkit.models.event import MediaContent
    
    await kit.process_inbound(
        InboundMessage(
            channel_id="ws-user",
            sender_id="user",
            content=MediaContent(url="https://example.com/chart.png", mime_type="image/png"),
        )
    )
    # AI sees the image and responds with analysis
    

    Tool Call Events

    AIChannel automatically broadcasts ephemeral TOOL_CALL_START and TOOL_CALL_END events when executing tools. Subscribe to these for UI indicators:

    await kit.subscribe_room("room-1", my_callback)
    
    # Callback receives:
    # TOOL_CALL_START: {tool_calls: [{id, name, arguments}], round, channel_id}
    # TOOL_CALL_END: {tool_calls: [{id, name, result}], round, channel_id, duration_ms}
    

    VoiceChannel handles real-time audio conversations with speech-to-text, text-to-speech, and an audio processing pipeline.

    Basic Voice Setup

    from roomkit import RoomKit, VoiceChannel
    from roomkit.voice.pipeline import AudioPipelineConfig, MockVADProvider, VADEvent, VADEventType
    from roomkit.voice.stt.mock import MockSTTProvider
    from roomkit.voice.tts.mock import MockTTSProvider
    from roomkit.voice.backends.mock import MockVoiceBackend
    
    kit = RoomKit()
    
    # Create providers
    backend = MockVoiceBackend()
    stt = MockSTTProvider(transcriptions=["Hello, how can I help?"])
    tts = MockTTSProvider()
    vad = MockVADProvider(events=[
        VADEvent(type=VADEventType.SPEECH_START),
        None,
        VADEvent(type=VADEventType.SPEECH_END, audio_bytes=b"audio"),
    ])
    
    # Create voice channel with pipeline
    voice = VoiceChannel(
        "voice-agent",
        stt=stt,
        tts=tts,
        backend=backend,
        pipeline=AudioPipelineConfig(vad=vad),
    )
    kit.register_channel(voice)
    

    Joining a Voice Session

    # Create room and attach voice channel
    await kit.create_room(room_id="call-room")
    await kit.attach_channel("call-room", "voice-agent")
    
    # Join a participant to the voice session
    session = await kit.join(
        room_id="call-room",
        channel_id="voice-agent",
        participant_id="caller-1",
    )
    
    # Leave when done
    await kit.leave(session)
    

    STT Providers

    ProviderClassConfigExtra
    DeepgramDeepgramSTTProviderDeepgramConfigroomkit[deepgram]
    SherpaOnnxSherpaOnnxSTTProviderSherpaOnnxSTTConfigroomkit[sherpa-onnx]
    GradiumGradiumSTTProviderGradiumSTTConfigroomkit[gradium]
    Qwen3 ASRQwen3ASRProviderQwen3ASRConfigroomkit[qwen-asr]
    MockMockSTTProvider—built-in

    Use lazy loaders to avoid import-time dependency checks:

    from roomkit.voice import get_deepgram_provider, get_deepgram_config
    
    DeepgramSTTProvider = get_deepgram_provider()
    DeepgramConfig = get_deepgram_config()
    
    stt = DeepgramSTTProvider(DeepgramConfig(
        api_key="...",
        model="nova-2",
        language="en",
    ))
    

    TTS Providers

    ProviderClassConfigExtra
    ElevenLabsElevenLabsTTSProviderElevenLabsConfigroomkit[elevenlabs]
    SherpaOnnxSherpaOnnxTTSProviderSherpaOnnxTTSConfigroomkit[sherpa-onnx]
    GradiumGradiumTTSProviderGradiumTTSConfigroomkit[gradium]
    Qwen3Qwen3TTSProviderQwen3TTSConfigroomkit[qwen-tts]
    NeuTTSNeuTTSProviderNeuTTSConfigroomkit[neutts]
    Grok TTSGrokTTSProviderGrokTTSConfigxAI
    MockMockTTSProvider—built-in
    from roomkit.voice import get_elevenlabs_provider, get_elevenlabs_config
    
    ElevenLabsTTSProvider = get_elevenlabs_provider()
    ElevenLabsConfig = get_elevenlabs_config()
    
    tts = ElevenLabsTTSProvider(ElevenLabsConfig(
        api_key="...",
        voice_id="21m00Tcm4TlvDq8ikWAM",
        model="eleven_turbo_v2",
    ))
    

    Voice Backends

    Backends handle audio transport between the framework and participants:

    BackendClassExtraUse Case
    Local mic/speakerLocalAudioBackendroomkit[local-audio]Development/testing
    FastRTC (WebRTC)FastRTCVoiceBackendroomkit[fastrtc]Browser-based voice
    RTPRTPVoiceBackendroomkit[rtp]VoIP integration
    SIPSIPVoiceBackendroomkit[sip]Telephony
    WebTransportWebTransportBackendroomkit[webtransport]Low-latency web
    MockMockVoiceBackendbuilt-inTesting
    from roomkit.voice import get_local_audio_backend
    
    LocalAudioBackend = get_local_audio_backend()
    backend = LocalAudioBackend(sample_rate=16000, channels=1)
    

    Interruption Handling

    Four strategies for handling user speech during TTS playback:

    from roomkit.voice.interruption import InterruptionConfig, InterruptionStrategy
    
    voice = VoiceChannel(
        "voice",
        stt=stt,
        tts=tts,
        backend=backend,
        pipeline=AudioPipelineConfig(vad=vad),
        interruption=InterruptionConfig(
            strategy=InterruptionStrategy.CONFIRMED,
            min_speech_ms=300,  # Wait 300ms of sustained speech before interrupting
        ),
    )
    
    StrategyBehavior
    IMMEDIATEInterrupt on any detected speech
    CONFIRMEDWait for sustained speech (min_speech_ms). Default.
    SEMANTICUse BackchannelDetector to ignore "uh-huh", "yeah"
    DISABLEDNever interrupt TTS playback

    Voice Greeting

    Send a greeting when a session starts:

    await kit.send_greeting(
        room_id="call-room",
        channel_id="voice-agent",
        greeting="Welcome! How can I help you today?",
        session=session,
    )
    

    Or configure on the Agent:

    from roomkit import Agent
    
    agent = Agent(
        "voice-agent",
        provider=provider,
        greeting="Welcome! How can I help you today?",
        stt=stt,
        tts=tts,
        backend=backend,
        pipeline=AudioPipelineConfig(vad=vad),
    )
    

    Voice Hooks

    from roomkit import HookTrigger, HookExecution
    
    @kit.hook(HookTrigger.ON_SPEECH_START, execution=HookExecution.ASYNC)
    async def on_speech(event, ctx):
        print("User started speaking")
    
    @kit.hook(HookTrigger.ON_TRANSCRIPTION, execution=HookExecution.ASYNC)
    async def on_transcription(event, ctx):
        print(f"Transcription: {event.content.body}")
    
    @kit.hook(HookTrigger.BEFORE_TTS)
    async def before_tts(event, ctx):
        # Can modify or block TTS text
        return HookResult.allow()
    
    @kit.hook(HookTrigger.ON_BARGE_IN, execution=HookExecution.ASYNC)
    async def on_barge_in(event, ctx):
        print("User interrupted the AI")
    

    Audio Bridging

    Bridge audio between sessions for human-to-human voice calls:

    voice = VoiceChannel("voice", backend=backend, bridge=True)
    
    # With bridge + STT for live transcription
    voice = VoiceChannel("voice", stt=stt, backend=backend, bridge=True)
    

    Audio bridge supports N-party calls with mixing and cross-rate resampling.

    DTMF

    Send and detect DTMF tones:

    # Send DTMF
    await voice.send_dtmf(session, digit="1", duration_ms=160)
    
    # Detect DTMF via hook
    @kit.hook(HookTrigger.ON_DTMF, execution=HookExecution.ASYNC)
    async def on_dtmf(event, ctx):
        print(f"DTMF digit: {event.data['digit']}")
    

    The audio pipeline sits between the voice backend and STT/TTS, processing audio through pluggable stages.

    Pipeline Architecture

    Inbound:   Backend -> [Resampler] -> [Recorder] -> [AEC] -> [AGC] -> [Denoiser] -> VAD -> [Diarization] + [DTMF]
    Outbound:  TTS -> [PostProcessors] -> [Recorder] -> AEC.feed_reference -> [Resampler] -> Backend
    

    All stages are optional except VAD (required for speech detection). Stages in brackets are skipped if not configured.

    Pipeline Configuration

    from roomkit import VoiceChannel
    from roomkit.voice.pipeline import AudioPipelineConfig, VADConfig
    
    pipeline = AudioPipelineConfig(
        resampler=my_resampler,           # Sample rate conversion
        vad=my_vad_provider,              # Voice activity detection (required)
        denoiser=my_denoiser,             # Background noise removal
        diarization=my_diarizer,          # Speaker identification
        aec=my_aec,                       # Echo cancellation
        agc=my_agc,                       # Automatic gain control
        dtmf=my_dtmf_detector,            # DTMF tone detection
        recorder=my_recorder,             # Audio recording
        recording_config=my_rec_config,   # Recording settings
        turn_detector=my_turn_detector,   # Turn-taking detection
        vad_config=VADConfig(
            silence_threshold_ms=500,     # Silence before speech_end
        ),
    )
    
    voice = VoiceChannel(
        "voice",
        stt=stt,
        tts=tts,
        backend=backend,
        pipeline=pipeline,
    )
    

    Capability-Aware Skipping

    AEC and AGC stages automatically skip when the backend declares native capabilities:

    from roomkit.voice.base import VoiceCapability
    
    # If backend has NATIVE_AEC, pipeline skips the AEC stage
    # If backend has NATIVE_AGC, pipeline skips the AGC stage
    

    Pipeline Stages Reference

    Resampler

    Converts audio between sample rates (e.g., 8kHz SIP to 16kHz STT).

    from roomkit.voice.pipeline.resampler import LinearResamplerProvider
    
    resampler = LinearResamplerProvider()
    # Or use SincResampler for higher quality
    

    VAD (Voice Activity Detection)

    Detects speech start/end events. Required for the pipeline.

    from roomkit.voice.pipeline import MockVADProvider, VADEvent, VADEventType, VADConfig
    
    # Mock for testing
    vad = MockVADProvider(events=[
        VADEvent(type=VADEventType.SPEECH_START),
        None,  # No event for this frame
        VADEvent(type=VADEventType.SPEECH_END, audio_bytes=b"speech-data"),
    ])
    
    # Production: SherpaOnnx VAD (local, offline)
    from roomkit.voice import get_sherpa_onnx_vad_provider, get_sherpa_onnx_vad_config
    
    SherpaVAD = get_sherpa_onnx_vad_provider()
    SherpaVADConfig = get_sherpa_onnx_vad_config()
    vad = SherpaVAD(SherpaVADConfig(threshold=0.5))
    

    AEC (Acoustic Echo Cancellation)

    Removes echo from the microphone signal caused by speaker output.

    # Speex AEC
    from roomkit.voice import get_speex_aec_provider
    
    SpeexAEC = get_speex_aec_provider()
    aec = SpeexAEC(sample_rate=16000, frame_size=160, filter_length=1024)
    
    # WebRTC AEC
    # pip install roomkit[webrtc-aec]
    

    The pipeline feeds TTS audio as reference to the AEC via process_outbound().

    AGC (Automatic Gain Control)

    Normalizes audio volume levels.

    from roomkit.voice.pipeline.agc import AGCConfig
    from roomkit.voice.pipeline.agc.mock import MockAGCProvider
    
    agc = MockAGCProvider()
    

    Denoiser

    Removes background noise from audio.

    # RNNoise (local, CPU-based)
    from roomkit.voice import get_rnnoise_denoiser_provider
    
    RNNoise = get_rnnoise_denoiser_provider()
    denoiser = RNNoise()
    
    # ai|coustics Quail (cloud API)
    # pip install roomkit[aicoustics]
    
    # SherpaOnnx denoiser (local ONNX model)
    from roomkit.voice import get_sherpa_onnx_denoiser_provider, get_sherpa_onnx_denoiser_config
    
    SherpaDenoiser = get_sherpa_onnx_denoiser_provider()
    SherpaDenoiserConfig = get_sherpa_onnx_denoiser_config()
    denoiser = SherpaDenoiser(SherpaDenoiserConfig())
    

    Diarization

    Identifies different speakers in multi-speaker audio.

    from roomkit.voice.pipeline.diarization.mock import MockDiarizationProvider
    
    diarizer = MockDiarizationProvider()
    

    DTMF Detection

    Detects dual-tone multi-frequency signals (phone keypad tones).

    from roomkit.voice.pipeline.dtmf.mock import MockDTMFDetector
    
    dtmf = MockDTMFDetector()
    

    DTMF runs in parallel with other pipeline stages (before AEC/AGC/denoiser).

    Audio Recorder

    Records inbound and outbound audio.

    from roomkit.voice.pipeline.recorder.mock import MockAudioRecorder
    from roomkit.voice.pipeline.recorder import RecordingConfig
    
    recorder = MockAudioRecorder()
    config = RecordingConfig(
        format="wav",
        sample_rate=16000,
        channels=1,
    )
    

    Turn Detector

    Determines when a speaker's turn is complete (post-STT):

    from roomkit.voice.pipeline.turn.mock import MockTurnDetector
    
    turn = MockTurnDetector()
    
    # Production: Smart turn detection (ML-based)
    # pip install roomkit[smart-turn]
    

    Turn detection accumulates transcription fragments until is_complete=True, then routes the combined text.

    Backchannel Detector

    Classifies short utterances as backchannel (e.g., "uh-huh", "yeah") to prevent false interruptions:

    from roomkit.voice.pipeline.backchannel.mock import MockBackchannelDetector
    
    bc = MockBackchannelDetector()
    

    Used with InterruptionStrategy.SEMANTIC.

    Post-Processors

    Custom audio transformations on outbound TTS audio:

    from roomkit.voice.pipeline.postprocessor.base import AudioPostProcessor
    

    Interruption Handling

    The InterruptionHandler manages what happens when the user speaks during TTS playback:

    from roomkit.voice.interruption import InterruptionConfig, InterruptionStrategy
    
    config = InterruptionConfig(
        strategy=InterruptionStrategy.CONFIRMED,
        min_speech_ms=300,
    )
    
    StrategyWhen to Use
    IMMEDIATEFast response, accept false positives
    CONFIRMEDBalanced — waits for sustained speech
    SEMANTICIgnore backchannel ("uh-huh") using BackchannelDetector
    DISABLEDNever interrupt (e.g., announcements)

    AudioFrame

    Inbound audio is represented as AudioFrame:

    from roomkit.voice.audio_frame import AudioFrame
    
    frame = AudioFrame(
        data=b"\x00" * 320,   # Raw PCM bytes
        sample_rate=16000,     # Hz
        channels=1,            # Mono
        sample_width=2,        # 16-bit PCM
        timestamp_ms=0.0,
    )
    

    Pipeline stages annotate frame.metadata as they process: denoiser, vad, aec, agc, diarization, dtmf keys.

    Mock Providers for Testing

    Every pipeline stage has a mock provider with pre-configured event sequences:

    from roomkit.voice.pipeline import (
        MockVADProvider,
        MockDenoiserProvider,
        MockDiarizationProvider,
        MockAGCProvider,
        MockAECProvider,
        MockDTMFDetector,
        MockAudioRecorder,
        MockTurnDetector,
        MockBackchannelDetector,
    )
    

    Example with mock VAD:

    from roomkit.voice.pipeline import MockVADProvider, VADEvent, VADEventType
    
    vad = MockVADProvider(events=[
        VADEvent(type=VADEventType.SPEECH_START),
        None,
        VADEvent(type=VADEventType.SPEECH_END, audio_bytes=b"speech"),
    ])
    

    RealtimeVoiceChannel connects to speech-to-speech AI models that handle audio directly — no separate STT/TTS needed. The AI model receives audio and responds with audio.

    Basic Setup

    from roomkit import RoomKit, ChannelCategory
    from roomkit.channels.realtime_voice import RealtimeVoiceChannel
    
    kit = RoomKit()
    
    # Gemini Live example
    from roomkit.voice import get_gemini_live_provider
    
    GeminiLiveProvider = get_gemini_live_provider()
    
    realtime = RealtimeVoiceChannel(
        "realtime-voice",
        provider=GeminiLiveProvider(
            api_key="...",
            model="gemini-2.0-flash-live-001",
        ),
        system_prompt="You are a helpful voice assistant. Keep responses brief.",
    )
    kit.register_channel(realtime)
    
    await kit.create_room(room_id="voice-room")
    await kit.attach_channel("voice-room", "realtime-voice", category=ChannelCategory.INTELLIGENCE)
    

    Providers

    ProviderClassExtraDescription
    Google Gemini LiveGeminiLiveProviderroomkit[realtime-gemini]Gemini 2.0 speech-to-speech
    OpenAI RealtimeOpenAIRealtimeProviderroomkit[realtime-openai]GPT-4o realtime audio
    xAI GrokXAIRealtimeProvider—Grok speech-to-speech
    MockMockRealtimeProviderbuilt-inTesting
    # OpenAI Realtime
    from roomkit.voice import get_openai_realtime_provider
    
    OpenAIRealtime = get_openai_realtime_provider()
    provider = OpenAIRealtime(api_key="sk-...", model="gpt-4o-realtime-preview")
    
    # xAI Grok
    from roomkit.voice import get_xai_realtime_provider, get_xai_realtime_config
    
    XAIRealtime = get_xai_realtime_provider()
    XAIConfig = get_xai_realtime_config()
    provider = XAIRealtime(XAIConfig(api_key="..."))
    

    Audio Transports

    Transports handle the client-side audio (browser or device):

    TransportClassExtraUse Case
    WebSocketWebSocketRealtimeTransportroomkit[websocket]Browser via WebSocket
    FastRTC (WebRTC)FastRTCRealtimeTransportroomkit[fastrtc]Browser via WebRTC
    Local mic/speakerLocalAudioBackendroomkit[local-audio]Local development
    MockMockRealtimeTransportbuilt-inTesting
    from roomkit.voice import get_websocket_realtime_transport
    
    WSTransport = get_websocket_realtime_transport()
    transport = WSTransport(host="0.0.0.0", port=8765)
    
    realtime = RealtimeVoiceChannel(
        "realtime-voice",
        provider=provider,
        transport=transport,
        system_prompt="You are a helpful assistant.",
    )
    

    Joining a Session

    session = await kit.join(
        room_id="voice-room",
        channel_id="realtime-voice",
        participant_id="caller-1",
    )
    
    # Leave when done
    await kit.leave(session)
    

    Tool Calling

    Realtime voice channels support tool calling during conversations:

    from roomkit.channels.realtime_voice import RealtimeVoiceChannel, ToolHandler
    
    async def handle_tool(name: str, arguments: dict) -> str:
        if name == "get_weather":
            return '{"temperature": 22, "condition": "sunny"}'
        return '{"error": "unknown tool"}'
    
    realtime = RealtimeVoiceChannel(
        "realtime-voice",
        provider=provider,
        system_prompt="You help with weather. Use the get_weather tool.",
        tools=[
            {
                "name": "get_weather",
                "description": "Get weather for a city",
                "parameters": {
                    "type": "object",
                    "properties": {"city": {"type": "string"}},
                    "required": ["city"],
                },
            },
        ],
        tool_handler=handle_tool,
    )
    

    Text Injection

    Inject text into an active realtime session (appears as if the AI said it):

    session = await kit.join("room", "realtime-voice", participant_id="user")
    await realtime.inject_text(session, "Let me check that for you...")
    

    VAD Configuration

    Configure voice activity detection for realtime providers:

    realtime = RealtimeVoiceChannel(
        "realtime-voice",
        provider=provider,
        system_prompt="...",
        vad_config={
            "threshold": 0.5,
            "silence_duration_ms": 500,
        },
    )
    

    Session Resumption

    Some providers support session resumption after disconnection:

    # OpenAI Realtime supports session resumption
    provider = OpenAIRealtime(
        api_key="sk-...",
        model="gpt-4o-realtime-preview",
    )
    # Sessions automatically resume when possible
    

    Hooks

    from roomkit import HookTrigger, HookExecution
    
    @kit.hook(HookTrigger.ON_TOOL_CALL, execution=HookExecution.ASYNC)
    async def on_tool(event, ctx):
        print(f"Realtime tool call: {event}")
    
    @kit.hook(HookTrigger.ON_REALTIME_TEXT_INJECTED, execution=HookExecution.ASYNC)
    async def on_inject(event, ctx):
        print(f"Text injected: {event}")
    

    RoomKit provides four declarative orchestration strategies for multi-agent workflows. Pass a strategy to RoomKit(orchestration=...) or create_room(orchestration=...) — agents, routing, handoff tools, and conversation state are wired automatically.

    Strategies

    Pipeline

    Linear agent chain: triage -> handler -> resolver. Each agent can only hand off to the next in sequence.

    from roomkit import Agent, Pipeline, RoomKit, WebSocketChannel
    from roomkit.providers.ai.mock import MockAIProvider
    
    triage = Agent(
        "triage",
        provider=MockAIProvider(responses=["Transferring you..."]),
        role="Triage agent",
        description="Routes requests to the right specialist",
        system_prompt="You triage incoming requests.",
    )
    handler = Agent(
        "handler",
        provider=MockAIProvider(responses=["Let me help with that."]),
        role="Request handler",
        description="Handles customer requests",
        system_prompt="You handle requests.",
    )
    resolver = Agent(
        "resolver",
        provider=MockAIProvider(responses=["All done!"]),
        role="Resolution specialist",
        description="Confirms resolution",
        system_prompt="You resolve and close requests.",
    )
    
    kit = RoomKit(orchestration=Pipeline(agents=[triage, handler, resolver]))
    

    Swarm

    Every agent can hand off to every other agent. Bidirectional routing.

    from roomkit import Swarm
    
    kit = RoomKit(orchestration=Swarm(agents=[billing, shipping, returns]))
    

    Supervisor

    A supervisor agent delegates tasks to worker agents in child rooms:

    from roomkit import Supervisor
    
    kit = RoomKit(orchestration=Supervisor(
        supervisor=manager_agent,
        workers=[researcher, writer, reviewer],
    ))
    

    Loop

    Producer/reviewer cycle. The reviewer has an approve_output tool to break the loop.

    from roomkit import Loop
    
    kit = RoomKit(orchestration=Loop(
        agent=writer_agent,
        reviewer=editor_agent,
        max_iterations=3,
    ))
    

    Using Orchestration

    from roomkit import InboundMessage, TextContent, WebSocketChannel
    
    # Register transport channel
    ws = WebSocketChannel("ws-user")
    kit.register_channel(ws)
    
    # Create room — orchestration auto-registers agents, creates router, sets initial state
    await kit.create_room(room_id="support")
    await kit.attach_channel("support", "ws-user")
    
    # Messages are automatically routed to the active agent
    await kit.process_inbound(
        InboundMessage(
            channel_id="ws-user",
            sender_id="user",
            content=TextContent(body="I need help with billing."),
        )
    )
    

    Handoff Protocol

    Agents hand off conversations by calling the handoff_conversation tool (auto-injected by orchestration strategies):

    # The AI calls this tool automatically:
    # handoff_conversation(target="handler", reason="Billing issue", summary="User needs invoice help")
    

    The handoff:

    1. Updates ConversationState in room metadata (phase, active agent, handoff count)
    2. Mutes the outgoing agent, unmutes the incoming agent
    3. Emits a system event visible to the new agent
    4. Records the transition in phase_history

    Conversation State

    ConversationState tracks conversation progress within a room. It's stored in Room.metadata["_conversation_state"] and persists across all message turns. Orchestration strategies create and update it automatically, but you can also read and modify it directly.

    ConversationState Fields

    FieldTypeDefaultDescription
    phasestr"intake"Current conversation phase. Can be any string — not restricted to built-in phases.
    active_agent_idstr | NoneNoneChannel ID of the currently active agent.
    previous_agent_idstr | NoneNoneAgent that was active before the last transition.
    handoff_countint0Total number of agent handoffs.
    phase_started_atdatetimenowTimestamp when the current phase started.
    phase_historylist[PhaseTransition][]Immutable audit trail of all transitions.
    contextdict[str, Any]{}Arbitrary user data — store any custom key-value pairs across turns.

    Built-in phase constants (use any string — these are convenience defaults): ConversationPhase.INTAKE, QUALIFICATION, HANDLING, ESCALATION, RESOLUTION, FOLLOWUP.

    Reading State

    from roomkit.orchestration.state import get_conversation_state
    
    room = await kit.get_room("support")
    state = get_conversation_state(room)
    
    print(state.phase)            # "handling"
    print(state.active_agent_id)  # "billing-agent"
    print(state.handoff_count)    # 2
    print(state.context)          # {"customer_tier": "premium", "issue_type": "refund"}
    
    # Phase transition history (audit trail)
    for t in state.phase_history:
        print(f"{t.from_phase} -> {t.to_phase} by {t.from_agent} -> {t.to_agent} ({t.reason})")
    

    Persisting Custom Data Across Turns

    Use state.context to store arbitrary data that survives across conversation turns. Update state via set_conversation_state() + store.update_room():

    from roomkit.orchestration.state import get_conversation_state, set_conversation_state
    
    # Read current state
    room = await kit.get_room("support")
    state = get_conversation_state(room)
    
    # Store custom data in context (immutable update pattern)
    updated_state = state.model_copy(update={
        "context": {
            **state.context,
            "customer_tier": "premium",
            "issue_type": "refund",
            "attempts": state.context.get("attempts", 0) + 1,
        }
    })
    
    # Persist back to room
    updated_room = set_conversation_state(room, updated_state)
    await kit.store.update_room(updated_room)
    

    Retrieving Custom Data on Later Turns

    # In a hook or handler on a subsequent turn:
    room = await kit.get_room("support")
    state = get_conversation_state(room)
    
    tier = state.context.get("customer_tier", "standard")
    attempts = state.context.get("attempts", 0)
    

    Programmatic Phase Transitions

    Use state.transition() to change phase and record an audit entry. It returns a new immutable state:

    from roomkit.orchestration.state import get_conversation_state, set_conversation_state
    
    room = await kit.get_room("support")
    state = get_conversation_state(room)
    
    # Transition to a new phase with reason and agent
    new_state = state.transition(
        to_phase="escalation",
        to_agent="supervisor-agent",
        reason="Customer requested manager",
        metadata={"escalation_priority": "high"},
    )
    # new_state.phase == "escalation"
    # new_state.active_agent_id == "supervisor-agent"
    # new_state.previous_agent_id == <old agent>
    # new_state.handoff_count incremented if agent changed
    
    # Persist
    updated_room = set_conversation_state(room, new_state)
    await kit.store.update_room(updated_room)
    

    Using State in Hooks

    A common pattern — persist data via a BEFORE_BROADCAST hook:

    from roomkit import HookTrigger, HookResult, RoomEvent, RoomContext, TextContent
    from roomkit.orchestration.state import get_conversation_state, set_conversation_state
    
    @kit.hook(HookTrigger.BEFORE_BROADCAST)
    async def track_sentiment(event: RoomEvent, ctx: RoomContext) -> HookResult:
        if not isinstance(event.content, TextContent):
            return HookResult.allow()
    
        state = get_conversation_state(ctx.room)
    
        # Update context with per-turn data
        updated_state = state.model_copy(update={
            "context": {
                **state.context,
                "last_message_length": len(event.content.body),
                "message_count": state.context.get("message_count", 0) + 1,
            }
        })
        updated_room = set_conversation_state(ctx.room, updated_state)
        await kit.store.update_room(updated_room)
    
        return HookResult.allow()
    

    PhaseTransition Audit Record

    Each transition creates an immutable PhaseTransition:

    FieldTypeDescription
    from_phasestrPrevious phase
    to_phasestrNew phase
    from_agentstr | NonePrevious agent channel ID
    to_agentstr | NoneNew agent channel ID
    reasonstrWhy the transition occurred
    timestampdatetimeWhen the transition occurred
    metadatadict[str, Any]Arbitrary metadata for this transition

    Conversation Router (Advanced)

    ConversationRouter dynamically routes incoming messages to different agents based on conversation state, message content, origin channel, or custom logic. It works as a BEFORE_BROADCAST hook that stamps routing metadata on events.

    RoutingConditions Reference

    All conditions are ANDed — every non-None field must match for the rule to fire:

    FieldTypeDescription
    phasesset[str] | NoneMatch when state.phase is in this set
    channel_typesset[ChannelType] | NoneMatch when sender's channel type is in this set
    intentsset[str] | NoneMatch when event.metadata["intent"] is in this set
    source_channel_idsset[str] | NoneMatch when sender's channel ID is in this set
    customCallable | NoneCustom function (event, context, state) -> bool for arbitrary logic

    Routing by Phase

    from roomkit.orchestration.router import ConversationRouter, RoutingRule, RoutingConditions
    
    router = ConversationRouter(
        rules=[
            RoutingRule(
                agent_id="billing-agent",
                conditions=RoutingConditions(phases={"billing"}),
                priority=0,
            ),
            RoutingRule(
                agent_id="shipping-agent",
                conditions=RoutingConditions(phases={"shipping"}),
                priority=0,
            ),
        ],
        default_agent_id="triage-agent",
    )
    

    Routing by Channel Type (Origin)

    Route messages to different agents based on where they came from:

    from roomkit.models.enums import ChannelType
    
    router = ConversationRouter(
        rules=[
            RoutingRule(
                agent_id="voice-specialist",
                conditions=RoutingConditions(channel_types={ChannelType.VOICE}),
            ),
            RoutingRule(
                agent_id="sms-agent",
                conditions=RoutingConditions(channel_types={ChannelType.SMS, ChannelType.WHATSAPP}),
            ),
        ],
        default_agent_id="general-agent",
    )
    

    Routing by Intent (Content-Based)

    Set event.metadata["intent"] via a classification hook, then route by intent:

    @kit.hook(HookTrigger.BEFORE_BROADCAST, priority=-200)  # Run before router
    async def classify_intent(event: RoomEvent, ctx: RoomContext) -> HookResult:
        if isinstance(event.content, TextContent):
            body = event.content.body.lower()
            intent = "billing" if "invoice" in body or "charge" in body else "general"
            modified = event.model_copy(update={"metadata": {**(event.metadata or {}), "intent": intent}})
            return HookResult.modify(modified)
        return HookResult.allow()
    
    router = ConversationRouter(
        rules=[
            RoutingRule(
                agent_id="billing-agent",
                conditions=RoutingConditions(intents={"billing"}),
            ),
        ],
        default_agent_id="general-agent",
    )
    

    Routing with Custom Logic

    Use a custom callable for complex routing decisions, including content-based or state-based logic:

    def is_high_value_customer(event: RoomEvent, ctx: RoomContext, state: ConversationState) -> bool:
        return state.context.get("customer_tier") == "premium"
    
    def contains_urgency(event: RoomEvent, ctx: RoomContext, state: ConversationState) -> bool:
        if isinstance(event.content, TextContent):
            return any(word in event.content.body.lower() for word in ["urgent", "emergency", "asap"])
        return False
    
    router = ConversationRouter(
        rules=[
            # Premium customers with urgent messages -> senior agent
            RoutingRule(
                agent_id="senior-agent",
                conditions=RoutingConditions(custom=is_high_value_customer),
                priority=-1,  # Check first
            ),
            # Any urgent message -> escalation agent
            RoutingRule(
                agent_id="escalation-agent",
                conditions=RoutingConditions(custom=contains_urgency),
                priority=0,
            ),
        ],
        default_agent_id="general-agent",
        supervisor_id="supervisor-agent",  # Always observes all conversations
    )
    

    Combined Conditions

    Combine multiple conditions (all are ANDed):

    # Only route SMS messages during billing phase to the billing specialist
    RoutingRule(
        agent_id="sms-billing-agent",
        conditions=RoutingConditions(
            phases={"billing"},
            channel_types={ChannelType.SMS},
        ),
        priority=0,
    )
    

    Installing the Router

    # Option 1: Manual hook installation
    kit.hook(
        HookTrigger.BEFORE_BROADCAST,
        execution=HookExecution.SYNC,
        priority=-100,
    )(router.as_hook())
    
    # Option 2: Use install() which also sets up handoff tools
    handler = router.install(kit, agents=[billing_agent, shipping_agent, triage_agent])
    

    Routing Selection Priority

    The router evaluates in this order:

    1. Agent affinity — if state.active_agent_id is set and the agent is still attached, stick with it
    2. Rules — evaluate RoutingRule list in ascending priority order; first match wins
    3. Fallback — return default_agent_id
    4. Loop prevention — events FROM intelligence channels are never routed (returns None)

    Conversation Pipeline (Advanced)

    Define stages explicitly:

    from roomkit.orchestration.pipeline import ConversationPipeline, PipelineStage
    
    pipeline = ConversationPipeline(stages=[
        PipelineStage(phase="triage", agent_id="triage", next="handling"),
        PipelineStage(phase="handling", agent_id="handler", next="resolution"),
        PipelineStage(phase="resolution", agent_id="resolver"),
    ])
    

    Status Bus

    Agents can publish status updates for UI display:

    # Agents publish status via the status bus
    # Subscribe to updates
    async def on_status(update):
        print(f"Agent {update.agent_id}: {update.message} ({update.level})")
    
    await kit.status_bus.subscribe(on_status)
    

    Delegation

    Delegate tasks to background agents in child rooms:

    result = await kit.delegate(
        room_id="main-room",
        agent_id="researcher",
        task="Find the latest pricing for product X",
        context={"product": "X"},
    )
    

    Memory Providers

    Agents use memory providers to maintain conversation context across handoffs:

    from roomkit.memory.sliding_window import SlidingWindowMemory
    from roomkit.orchestration.handoff import HandoffMemoryProvider
    
    agent = Agent(
        "agent",
        provider=provider,
        memory=HandoffMemoryProvider(SlidingWindowMemory(max_events=50)),
    )
    

    HandoffMemoryProvider wraps any memory provider to inject handoff context (summary from the previous agent) into the conversation history.

    Transport providers handle sending and receiving messages over external protocols. Each provider implements a channel-specific ABC.

    SMS

    Twilio

    from roomkit import RoomKit, SMSChannel
    from roomkit.providers.twilio.sms import TwilioSMSProvider
    from roomkit.providers.twilio.config import TwilioConfig
    
    sms = SMSChannel("sms-twilio", provider=TwilioSMSProvider(TwilioConfig(
        account_sid="AC...",
        auth_token="...",
        from_number="+15551234567",
    )))
    
    kit = RoomKit()
    kit.register_channel(sms)
    

    Telnyx

    from roomkit.providers.telnyx.sms import TelnyxSMSProvider
    from roomkit.providers.telnyx.config import TelnyxConfig
    
    sms = SMSChannel("sms-telnyx", provider=TelnyxSMSProvider(TelnyxConfig(
        api_key="KEY...",
        from_number="+15551234567",
    )))
    

    Sinch

    from roomkit.providers.sinch.sms import SinchSMSProvider
    from roomkit.providers.sinch.config import SinchConfig
    
    sms = SMSChannel("sms-sinch", provider=SinchSMSProvider(SinchConfig(
        service_plan_id="...",
        api_token="...",
        from_number="+15551234567",
    )))
    

    VoiceMeUp

    from roomkit.providers.voicemeup.sms import VoiceMeUpSMSProvider
    from roomkit.providers.voicemeup.config import VoiceMeUpConfig
    
    sms = SMSChannel("sms-vmu", provider=VoiceMeUpSMSProvider(VoiceMeUpConfig(
        username="...",
        auth_token="...",
        from_number="+15551234567",
    )))
    

    Webhook Parsing

    Each SMS provider has a webhook parser:

    from roomkit.providers.twilio.sms import parse_twilio_webhook
    from roomkit.providers.telnyx.sms import parse_telnyx_webhook
    from roomkit.providers.sinch.sms import parse_sinch_webhook
    from roomkit.providers.voicemeup.sms import parse_voicemeup_webhook
    
    # Or use the universal webhook parser
    message = await kit.process_webhook(meta=request_data, channel_id="sms-twilio")
    

    RCS

    from roomkit import RCSChannel
    from roomkit.providers.twilio.rcs import TwilioRCSProvider, TwilioRCSConfig
    
    rcs = RCSChannel("rcs-main", provider=TwilioRCSProvider(TwilioRCSConfig(
        account_sid="AC...",
        auth_token="...",
        messaging_service_sid="MG...",  # Required for RCS (must be RCS-enabled)
    )))
    

    Also available via Telnyx: TelnyxRCSProvider, TelnyxRCSConfig.

    Email

    Elastic Email

    from roomkit import EmailChannel
    from roomkit.providers.elasticemail.email import ElasticEmailProvider
    from roomkit.providers.elasticemail.config import ElasticEmailConfig
    
    email = EmailChannel("email-main", provider=ElasticEmailProvider(ElasticEmailConfig(
        api_key="...",
        from_email="support@example.com",
        from_name="Support Team",
    )))
    

    SendGrid

    from roomkit.providers.sendgrid.email import SendGridEmailProvider
    from roomkit.providers.sendgrid.config import SendGridConfig
    
    email = EmailChannel("email-sg", provider=SendGridEmailProvider(SendGridConfig(
        api_key="SG...",
        from_email="support@example.com",
    )))
    

    WhatsApp

    Business API (Cloud)

    from roomkit import WhatsAppChannel
    from roomkit.providers.whatsapp.base import WhatsAppProvider
    
    whatsapp = WhatsAppChannel("wa-business", provider=WhatsAppProvider(
        access_token="...",
        phone_number_id="...",
    ))
    

    Personal (neonize)

    from roomkit import WhatsAppPersonalChannel
    from roomkit.providers.whatsapp.personal import WhatsAppPersonalProvider
    
    whatsapp = WhatsAppPersonalChannel("wa-personal", provider=WhatsAppPersonalProvider())
    

    Requires pip install roomkit[whatsapp-personal]. Uses the neonize library for multidevice protocol with typing indicators, read receipts, and media handling.

    Facebook Messenger

    from roomkit import MessengerChannel
    from roomkit.providers.messenger.facebook import FacebookMessengerProvider
    from roomkit.providers.messenger.config import MessengerConfig
    
    messenger = MessengerChannel("messenger", provider=FacebookMessengerProvider(MessengerConfig(
        page_access_token="...",
        app_secret="...",
        verify_token="...",
    )))
    

    Webhook parser: parse_messenger_webhook(request_data, channel_id="messenger") — returns list[InboundMessage].

    Telegram

    from roomkit import TelegramChannel
    from roomkit.providers.telegram.bot import TelegramBotProvider
    from roomkit.providers.telegram.config import TelegramConfig
    
    telegram = TelegramChannel("telegram", provider=TelegramBotProvider(TelegramConfig(
        bot_token="123456:ABC-DEF...",
    )))
    

    Webhook parser: parse_telegram_webhook(request_data, channel_id="telegram") — returns list[InboundMessage].

    Microsoft Teams

    from roomkit import TeamsChannel
    from roomkit.providers.teams.bot_framework import BotFrameworkTeamsProvider
    from roomkit.providers.teams.config import TeamsConfig
    
    teams = TeamsChannel("teams", provider=BotFrameworkTeamsProvider(TeamsConfig(
        app_id="...",
        app_password="...",
    )))
    

    Features: proactive messaging, bot mention detection, reaction handling, conversation reference storage.

    from roomkit.providers.teams.webhook import parse_teams_webhook, is_bot_added
    
    # Parse incoming Teams activity
    activity = parse_teams_webhook(request_data)
    
    # Check if bot was added to a conversation
    if is_bot_added(activity):
        # Handle bot installation
        pass
    

    HTTP (Generic Webhook)

    from roomkit import HTTPChannel
    from roomkit.providers.http.provider import WebhookHTTPProvider
    from roomkit.providers.http.config import HTTPProviderConfig
    
    http = HTTPChannel("webhook", provider=WebhookHTTPProvider(HTTPProviderConfig(
        url="https://api.example.com/messages",
        headers={"Authorization": "Bearer ..."},
    )))
    

    WebSocket

    WebSocket channels don't use a provider — they handle connections directly:

    from roomkit import WebSocketChannel
    
    ws = WebSocketChannel("ws-client")
    
    # Register a connection
    ws.register_connection("conn-1", on_receive_callback)
    
    # In production, connect to the framework
    await kit.connect_websocket("ws-client", "conn-1", send_fn)
    await kit.disconnect_websocket("ws-client", "conn-1")
    

    Phone Number Utilities

    from roomkit.providers.sms.phone import is_valid_phone, normalize_phone
    
    is_valid_phone("+15551234567")   # True
    normalize_phone("555-123-4567")  # "+15551234567"
    

    Delivery Status Tracking

    Track delivery status for sent messages:

    from roomkit import DeliveryStatus
    
    @kit.on_delivery_status
    async def track_delivery(status: DeliveryStatus) -> None:
        if status.status == "failed":
            print(f"Message {status.message_id} failed: {status.error_message}")
    
    # Process status webhooks from providers
    await kit.process_delivery_status(status)
    

    Identity Resolution

    The identity pipeline maps external sender IDs to known participants. It runs as part of the inbound pipeline, after handle_inbound() and before hooks.

    How It Works

    Inbound message arrives with sender_id
      -> IdentityResolver.resolve(sender_id, channel_type)
      -> Returns IdentityResult with status:
         IDENTIFIED      -> participant_id stamped on event, processing continues
         AMBIGUOUS       -> ON_IDENTITY_AMBIGUOUS hook fires
         PENDING         -> ON_IDENTITY_AMBIGUOUS hook fires
         UNKNOWN         -> ON_IDENTITY_UNKNOWN hook fires
         REJECTED        -> ON_IDENTITY_UNKNOWN hook fires
    

    Identity Hooks

    from roomkit import RoomKit, HookTrigger
    from roomkit.models.identity import IdentityHookResult, Identity
    
    kit = RoomKit()
    
    @kit.identity_hook(HookTrigger.ON_IDENTITY_UNKNOWN)
    async def handle_unknown(event, ctx):
        # Option 1: Resolve to a known identity
        return IdentityHookResult.resolved(Identity(
            id="user-123",
            display_name="Alice",
        ))
    
        # Option 2: Challenge the sender to identify
        return IdentityHookResult.challenge(inject=InjectedEvent(
            content=TextContent(body="Please provide your account number."),
        ))
    
        # Option 3: Reject the message
        return IdentityHookResult.reject("Unknown sender")
    
        # Option 4: Keep as pending
        return IdentityHookResult.pending(candidates=[...])
    

    Custom Identity Resolver

    from roomkit.identity.base import IdentityResolver
    from roomkit.models.identity import IdentityResult, Identity
    from roomkit.models.enums import IdentificationStatus
    
    class DatabaseIdentityResolver(IdentityResolver):
        async def resolve(self, message: InboundMessage, context: RoomContext) -> IdentityResult:
            user = await db.find_by_phone(message.sender_id)
            if user:
                return IdentityResult(
                    status=IdentificationStatus.IDENTIFIED,
                    identity=Identity(id=user.id, display_name=user.name),
                )
            return IdentityResult(status=IdentificationStatus.UNKNOWN)
    
    kit = RoomKit(identity_resolver=DatabaseIdentityResolver())
    

    Manual Resolution

    # Resolve a pending participant to a known identity
    await kit.resolve_participant(
        room_id="room-1",
        participant_id="pending-123",
        identity_id="user-456",
    )
    

    Realtime Ephemeral Events

    Ephemeral events (typing, presence, reactions) are not stored in conversation history. They're delivered in real-time to subscribers.

    Publishing Events

    from roomkit import RoomKit
    
    kit = RoomKit()
    
    # Typing indicator
    await kit.publish_typing("room-1", "alice", is_typing=True)
    await kit.publish_typing("room-1", "alice", is_typing=False)
    
    # Presence
    await kit.publish_presence("room-1", "alice", "online")   # online/away/offline
    
    # Reaction
    await kit.publish_reaction("room-1", "alice", target_event_id="evt-123", emoji="thumbsup")
    
    # Read receipt
    await kit.publish_read_receipt("room-1", "alice", event_id="evt-123")
    
    # Tool call events (AIChannel publishes these automatically)
    from roomkit.realtime.base import EphemeralEventType
    
    await kit.publish_tool_call("room-1", "ai-agent", [
        {"id": "tc1", "name": "search", "arguments": {"q": "test"}}
    ], EphemeralEventType.TOOL_CALL_START)
    

    Subscribing to Events

    async def on_ephemeral(event):
        print(f"Ephemeral: {event.type} from {event.user_id}")
    
    sub_id = await kit.subscribe_room("room-1", on_ephemeral)
    
    # Unsubscribe later
    await kit.unsubscribe_room(sub_id)
    

    Event Types

    TypeDescription
    TYPING_STARTUser started typing
    TYPING_STOPUser stopped typing
    PRESENCE_ONLINEUser came online
    PRESENCE_AWAYUser went away
    PRESENCE_OFFLINEUser went offline
    READ_RECEIPTUser read a message
    REACTIONUser reacted to a message
    TOOL_CALL_STARTAI started executing a tool
    TOOL_CALL_ENDAI finished executing a tool
    CUSTOMCustom ephemeral event

    Read Tracking

    # Mark a specific event as read
    await kit.mark_read("room-1", "ws-user", "evt-123")
    
    # Mark all events as read
    await kit.mark_all_read("room-1", "ws-user")
    

    Custom Realtime Backend

    The default InMemoryRealtime works for single-process deployments. For distributed systems, implement RealtimeBackend:

    from roomkit.realtime.base import RealtimeBackend
    
    class RedisRealtimeBackend(RealtimeBackend):
        # Implement publish, subscribe, unsubscribe using Redis Pub/Sub
        ...
    
    kit = RoomKit(realtime=RedisRealtimeBackend())
    

    RoomKit provides built-in resilience patterns for production deployments: rate limiting, circuit breakers, retry with backoff, chain depth limits, and room lifecycle timers.

    Rate Limiting

    Apply rate limits per channel binding:

    from roomkit import RoomKit
    from roomkit.models.channel import RateLimit
    
    kit = RoomKit()
    
    await kit.attach_channel("room-1", "sms-out",
        rate_limit=RateLimit(max_per_second=1.0, max_per_minute=30.0),
    )
    

    Framework-level inbound rate limiting:

    from roomkit import RoomKit
    from roomkit.models.channel import RateLimit
    
    kit = RoomKit(inbound_rate_limit=RateLimit(max_per_second=10.0))
    

    Circuit Breaker

    Isolate provider failures with circuit breakers:

    from roomkit.core.circuit_breaker import CircuitBreaker
    
    cb = CircuitBreaker(failure_threshold=5, recovery_timeout=60.0)
    
    if cb.allow_request():
        try:
            result = await provider.send(event, to="+1234567890")
            cb.record_success()
        except Exception:
            cb.record_failure()  # Opens after 5 consecutive failures
    

    States: CLOSED (normal) -> OPEN (failing, fast-reject) -> HALF_OPEN (testing recovery).

    Retry with Backoff

    Retry failed operations with exponential backoff:

    from roomkit.models.channel import RetryPolicy
    from roomkit.core.retry import retry_with_backoff
    
    policy = RetryPolicy(
        max_retries=3,
        base_delay_seconds=1.0,
        max_delay_seconds=60.0,
    )
    
    result = await retry_with_backoff(flaky_function, policy)
    

    Apply per binding:

    await kit.attach_channel("room-1", "sms-out",
        retry_policy=RetryPolicy(max_retries=3, base_delay_seconds=1.0),
    )
    

    Chain Depth Limit

    Prevents infinite AI-to-AI loops. Default max depth is 5.

    kit = RoomKit(max_chain_depth=3)  # Stricter limit
    

    When an AI response triggers another AI response, chain depth increments. Processing stops when the limit is reached.

    Room Lifecycle Timers

    Auto-transition rooms based on inactivity:

    from roomkit import RoomKit
    
    kit = RoomKit()
    
    room = await kit.create_room(room_id="session-1")
    # Configure timers on the room:
    # inactive_after_seconds -> ACTIVE to PAUSED
    # closed_after_seconds -> to CLOSED
    
    # Check timers for one room
    room = await kit.check_room_timers("session-1")
    
    # Batch check all rooms (call periodically)
    transitioned = await kit.check_all_timers()
    

    Delivery Status Tracking

    Track whether messages were delivered to providers:

    from roomkit import DeliveryStatus
    
    @kit.on_delivery_status
    async def track(status: DeliveryStatus) -> None:
        if status.status == "failed":
            logger.error("Delivery failed: %s — %s", status.message_id, status.error_message)
        elif status.status == "delivered":
            logger.info("Delivered: %s", status.message_id)
    
    # Process status webhooks from providers
    await kit.process_delivery_status(status)
    

    Production Setup Example

    from roomkit import RoomKit
    from roomkit.models.channel import RateLimit, RetryPolicy
    from roomkit.store.postgres import PostgresStore
    
    kit = RoomKit(
        store=PostgresStore("postgresql://user:pass@localhost/roomkit"),
        max_chain_depth=5,
        inbound_rate_limit=RateLimit(max_per_second=50.0),
        process_timeout=30.0,
    )
    
    # Per-channel resilience
    await kit.attach_channel("room", "sms-out",
        rate_limit=RateLimit(max_per_second=1.0, max_per_minute=30.0),
        retry_policy=RetryPolicy(max_retries=3, base_delay_seconds=1.0),
    )
    

    Framework Events for Monitoring

    @kit.on("source_error")
    async def on_error(event):
        logger.error("Source error: %s", event.data["error"])
    
    @kit.on("voice_session_ended")
    async def on_voice_end(event):
        logger.info("Voice session ended: %s", event.data["session_id"])
    

    RoomKit uses the ConversationStore ABC for persistence. The default InMemoryStore works out of the box. For production, use PostgresStore.

    InMemoryStore (Default)

    from roomkit import RoomKit
    
    kit = RoomKit()  # Uses InMemoryStore automatically
    

    Data lives in Python dicts — fast for development, lost on restart.

    PostgresStore

    Install: pip install roomkit[postgres]

    from roomkit import RoomKit
    from roomkit.store.postgres import PostgresStore
    
    store = PostgresStore("postgresql://user:pass@localhost/roomkit")
    await store.init()  # Creates tables if they don't exist
    
    kit = RoomKit(store=store)
    

    Connection Pooling

    store = PostgresStore("postgresql://user:pass@localhost/roomkit")
    await store.init(min_size=5, max_size=20)  # Pool sizing via init()
    

    Schema

    PostgresStore creates 10 tables:

    TablePurpose
    roomsRoom records with status, metadata, timers
    eventsEvent timeline with sequential indexing
    participantsRoom participants with roles and status
    bindingsChannel-to-room bindings with config
    identitiesKnown identity records
    tasksAI-extracted tasks
    observationsAI-extracted observations
    delivery_statusMessage delivery tracking
    read_trackingPer-channel read positions
    telemetry_spansTelemetry span records

    Operations

    # Room operations
    room = await kit.create_room(room_id="persistent", metadata={"topic": "billing"})
    
    # Event storage and retrieval
    events = await kit.store.list_events("persistent", offset=0, limit=50)
    
    # Timeline query with filters
    timeline = await kit.get_timeline("persistent", offset=0, limit=50)
    
    # Participant management
    participants = await kit.store.list_participants("persistent")
    
    # Binding management
    bindings = await kit.store.list_bindings("persistent")
    

    Full Example

    from __future__ import annotations
    
    import asyncio
    import os
    
    from roomkit import InboundMessage, RoomKit, TextContent, WebSocketChannel
    from roomkit.store.postgres import PostgresStore
    
    
    async def main() -> None:
        store = PostgresStore(os.environ["DATABASE_URL"])
        await store.init()
    
        kit = RoomKit(store=store)
    
        ws = WebSocketChannel("ws-user")
        kit.register_channel(ws)
    
        room = await kit.create_room(room_id="support", metadata={"topic": "billing"})
        await kit.attach_channel("support", "ws-user")
    
        await kit.process_inbound(
            InboundMessage(
                channel_id="ws-user",
                sender_id="user",
                content=TextContent(body="I need help with my invoice"),
            )
        )
    
        # Data persists across restarts
        events = await kit.store.list_events("support")
        print(f"Stored {len(events)} events")
    
        await kit.close()
    
    
    if __name__ == "__main__":
        asyncio.run(main())
    

    Custom Store

    Implement ConversationStore for other backends (Redis, DynamoDB, etc.):

    from roomkit.store.base import ConversationStore
    from roomkit.models.room import Room
    
    class RedisStore(ConversationStore):
        async def create_room(self, room: Room) -> Room:
            await self.redis.set(f"room:{room.id}", room.model_dump_json())
            return room
    
        # Implement all abstract methods...
    
    kit = RoomKit(store=RedisStore())
    

    RoomKit provides mock implementations for every pluggable component: AI providers, voice backends, pipeline stages, identity resolvers, and telemetry.

    Test Setup

    pip install roomkit[dev]
    uv run pytest              # Run all tests
    uv run pytest tests/test_framework.py -v  # Specific file
    

    Configuration in pyproject.toml:

    [tool.pytest.ini_options]
    testpaths = ["tests"]
    asyncio_mode = "auto"  # No @pytest.mark.asyncio needed
    

    Basic Test Pattern

    from roomkit import RoomKit, InboundMessage, TextContent, WebSocketChannel
    
    
    class TestMyFeature:
        async def test_message_delivery(self) -> None:
            kit = RoomKit()
    
            ws = WebSocketChannel("ws-user")
            kit.register_channel(ws)
    
            inbox: list = []
            async def on_recv(_conn: str, event) -> None:
                inbox.append(event)
    
            ws.register_connection("conn", on_recv)
    
            await kit.create_room(room_id="test-room")
            await kit.attach_channel("test-room", "ws-user")
    
            result = await kit.process_inbound(
                InboundMessage(
                    channel_id="ws-user",
                    sender_id="user",
                    content=TextContent(body="Hello"),
                )
            )
    
            assert not result.blocked
            assert len(inbox) == 1
            assert inbox[0].content.body == "Hello"
    

    Mock AI Provider

    from roomkit.providers.ai.mock import MockAIProvider
    from roomkit.channels.ai import AIChannel
    from roomkit import ChannelCategory
    
    # Responds with pre-configured messages in order
    provider = MockAIProvider(responses=["Response 1", "Response 2"])
    
    ai = AIChannel("ai", provider=provider)
    kit.register_channel(ai)
    await kit.attach_channel("room", "ai", category=ChannelCategory.INTELLIGENCE)
    
    # After processing, check what the AI was called with:
    assert len(provider.calls) == 1
    last_call = provider.calls[-1]
    print(last_call.system_prompt)
    print(last_call.temperature)
    print([t.name for t in last_call.tools])
    

    Mock Voice Backend

    from roomkit.voice.backends.mock import MockVoiceBackend
    
    backend = MockVoiceBackend()
    
    # Simulate audio input
    backend.simulate_audio_received(session, audio_frame)
    

    Mock Pipeline Providers

    Every pipeline stage has a mock that accepts pre-configured event sequences:

    from roomkit.voice.pipeline import (
        MockVADProvider,
        VADEvent,
        VADEventType,
        MockDenoiserProvider,
        MockDiarizationProvider,
        MockAGCProvider,
        MockAECProvider,
        MockDTMFDetector,
        MockAudioRecorder,
        MockTurnDetector,
        MockBackchannelDetector,
    )
    
    # VAD with event sequence
    vad = MockVADProvider(events=[
        VADEvent(type=VADEventType.SPEECH_START),
        None,  # No event for this frame
        VADEvent(type=VADEventType.SPEECH_END, audio_bytes=b"speech"),
    ])
    
    # Other mocks
    denoiser = MockDenoiserProvider()
    diarizer = MockDiarizationProvider()
    agc = MockAGCProvider()
    aec = MockAECProvider()
    dtmf = MockDTMFDetector()
    recorder = MockAudioRecorder()
    turn = MockTurnDetector()
    backchannel = MockBackchannelDetector()
    

    Mock STT/TTS

    from roomkit.voice.stt.mock import MockSTTProvider
    from roomkit.voice.tts.mock import MockTTSProvider
    
    stt = MockSTTProvider(transcriptions=["Hello", "How are you?"])
    tts = MockTTSProvider()
    

    Mock Identity Resolver

    from roomkit.identity.mock import MockIdentityResolver
    
    resolver = MockIdentityResolver()
    kit = RoomKit(identity_resolver=resolver)
    

    Mock Realtime Provider

    from roomkit.voice.realtime.mock import MockRealtimeProvider, MockRealtimeTransport
    
    provider = MockRealtimeProvider()
    transport = MockRealtimeTransport()
    

    Testing Hooks

    from roomkit import HookTrigger, HookResult, HookExecution
    
    async def test_hook_blocks_message() -> None:
        kit = RoomKit()
        ws = WebSocketChannel("ws")
        kit.register_channel(ws)
    
        await kit.create_room(room_id="r")
        await kit.attach_channel("r", "ws")
    
        @kit.hook(HookTrigger.BEFORE_BROADCAST)
        async def blocker(event, ctx):
            return HookResult.block("blocked")
    
        result = await kit.process_inbound(
            InboundMessage(
                channel_id="ws",
                sender_id="user",
                content=TextContent(body="test"),
            )
        )
    
        assert result.blocked
        assert result.reason == "blocked"
    

    Testing Voice Pipeline

    from roomkit import VoiceChannel
    from roomkit.voice.pipeline import AudioPipelineConfig, MockVADProvider, VADEvent, VADEventType
    from roomkit.voice.stt.mock import MockSTTProvider
    from roomkit.voice.tts.mock import MockTTSProvider
    from roomkit.voice.backends.mock import MockVoiceBackend
    from roomkit.voice.audio_frame import AudioFrame
    
    async def test_voice_pipeline() -> None:
        kit = RoomKit()
    
        backend = MockVoiceBackend()
        stt = MockSTTProvider(transcriptions=["Hello"])
        tts = MockTTSProvider()
        vad = MockVADProvider(events=[
            VADEvent(type=VADEventType.SPEECH_START),
            None,
            VADEvent(type=VADEventType.SPEECH_END, audio_bytes=b"audio"),
        ])
    
        voice = VoiceChannel(
            "voice", stt=stt, tts=tts, backend=backend,
            pipeline=AudioPipelineConfig(vad=vad),
        )
        kit.register_channel(voice)
    
        await kit.create_room(room_id="call")
        await kit.attach_channel("call", "voice")
    
        # Simulate audio input
        frame = AudioFrame(data=b"\x00" * 320, sample_rate=16000)
        backend.simulate_audio_received(None, frame)
    

    Testing Orchestration

    from roomkit import Agent, Pipeline, RoomKit, WebSocketChannel
    from roomkit.providers.ai.mock import MockAIProvider
    
    async def test_pipeline_orchestration() -> None:
        agent1 = Agent("a1", provider=MockAIProvider(responses=["Transferring..."]))
        agent2 = Agent("a2", provider=MockAIProvider(responses=["Resolved!"]))
    
        kit = RoomKit(orchestration=Pipeline(agents=[agent1, agent2]))
    
        ws = WebSocketChannel("ws")
        kit.register_channel(ws)
    
        await kit.create_room(room_id="test")
        await kit.attach_channel("test", "ws")
    
        # First message goes to agent1
        result = await kit.process_inbound(
            InboundMessage(channel_id="ws", sender_id="user", content=TextContent(body="Help"))
        )
        assert not result.blocked
    

    Test Utilities

    # Pydantic model updates (immutable)
    modified = event.model_copy(update={"content": TextContent(body="new")})
    
    # Never mutate models directly — always use model_copy
    

    RoomKit exports 73 symbols from roomkit. Providers and voice types import from subpackages.

    Top-Level Imports (from roomkit import ...)

    Framework

    SymbolDescription
    RoomKitCentral orchestrator — rooms, channels, hooks, storage

    Channels

    SymbolDescription
    AgentAI agent with role, description, greeting, tools
    AIChannelIntelligence layer for AI responses
    AudioVideoChannelCombined audio + video channel
    ChannelBase class for all channels
    CLIChannelCLI/terminal interactive channel
    EmailChannelEmail transport channel factory
    HTTPChannelHTTP webhook transport channel factory
    MessengerChannelFacebook Messenger transport channel factory
    RCSChannelRCS transport channel factory
    RealtimeAudioVideoChannelRealtime speech-to-speech with video
    RealtimeVoiceChannelSpeech-to-speech AI channel
    SMSChannelSMS transport channel factory
    TeamsChannelMicrosoft Teams transport channel factory
    TelegramChannelTelegram Bot transport channel factory
    TransportChannelGeneric transport channel wrapper
    VideoChannelVideo channel with vision pipeline
    VoiceChannelReal-time audio with STT/TTS/pipeline
    WebSocketChannelWebSocket bidirectional channel
    WhatsAppChannelWhatsApp Business API channel factory
    WhatsAppPersonalChannelWhatsApp Personal (neonize) channel factory

    Enums

    SymbolDescription
    AccessChannel access levels: READ_WRITE, READ_ONLY, WRITE_ONLY, NONE
    ChannelCategoryTRANSPORT or INTELLIGENCE
    ChannelTypeSMS, EMAIL, WHATSAPP, VOICE, AI, WEBSOCKET, etc.
    EventStatusDELIVERED, BLOCKED, etc.
    EventTypeMESSAGE, SYSTEM, EDIT, DELETE, etc.
    HookExecutionSYNC or ASYNC
    HookTrigger40+ hook triggers (BEFORE_BROADCAST, AFTER_BROADCAST, etc.)
    RoomStatusACTIVE, PAUSED, CLOSED, ARCHIVED

    Orchestration

    SymbolDescription
    LoopProducer/reviewer cycle strategy
    OrchestrationABC for orchestration strategies
    PipelineLinear agent chain strategy
    SupervisorSupervisor delegates to workers strategy
    SwarmBidirectional handoff strategy

    Models

    SymbolDescription
    ChannelBindingBinding of a channel to a room
    ChannelCapabilitiesDeclared capabilities of a channel
    ChannelOutputOutput of a channel delivery
    DeliveryResultResult of delivering a message
    DeliveryStatusDelivery status from provider webhook
    EventSourceSource attribution for an event
    FrameworkEventLightweight framework lifecycle event
    HookResultResult from sync hooks: .allow(), .block(reason), .modify(event)
    InjectedEventEvent injected by a hook
    InboundMessageIncoming message from a provider
    InboundResultResult of processing an inbound message
    ParticipantParticipant data model
    ProviderResultResult from a provider operation
    RoomRoom data model
    RoomContextContext passed to hooks (room, bindings, participants, events)
    RoomEventCore event stored in the timeline
    RoomTimersTimer configuration for room inactivity
    SessionStartedEventEvent fired when a voice session starts
    TextContentPlain text content
    ToolBase class for tool definitions
    ToolCallCallbackCallback type for tool call events
    ToolCallEventTool call event model
    ToolHandlerTool handler type for realtime voice
    get_current_voice_sessionGet the current voice session from context

    Errors

    SymbolDescription
    RoomKitErrorBase exception
    RoomNotFoundErrorRoom does not exist
    ChannelNotFoundErrorChannel not attached to room
    ChannelNotRegisteredErrorChannel not registered with framework
    ParticipantNotFoundErrorParticipant not found in room
    IdentityNotFoundErrorIdentity not found
    SourceAlreadyAttachedErrorSource already attached
    SourceNotFoundErrorNo source attached
    VoiceBackendNotConfiguredErrorVoice backend not configured
    VoiceNotConfiguredErrorVoice (STT/TTS) not configured

    AI Documentation Helpers

    SymbolDescription
    get_llms_txt()Get llms.txt content
    get_llms_full_txt()Get llms-full.txt content (comprehensive)
    get_agents_md()Get AGENTS.md content
    get_ai_context()Get combined AI context

    RoomKit Constructor

    kit = RoomKit(
        store=InMemoryStore(),                 # ConversationStore implementation
        identity_resolver=None,                # IdentityResolver implementation
        identity_channel_types=None,           # Channel types to resolve identity for
        inbound_router=None,                   # InboundRoomRouter implementation
        lock_manager=None,                     # RoomLockManager implementation
        realtime=None,                         # RealtimeBackend implementation
        max_chain_depth=5,                     # AI-to-AI loop prevention
        identity_timeout=10.0,                 # Identity resolution timeout (seconds)
        process_timeout=30.0,                  # Inbound processing timeout (seconds)
        stt=None,                              # STTProvider for transcription
        tts=None,                              # TTSProvider for synthesis
        voice=None,                            # VoiceBackend for audio transport
        task_runner=None,                      # Background task runner
        delivery_strategy=None,                # Delivery strategy
        status_bus=None,                       # Status bus for orchestration
        telemetry=None,                        # TelemetryProvider
        inbound_rate_limit=None,               # Framework-level rate limit
        orchestration=None,                    # Orchestration strategy
    )
    

    Key RoomKit Methods

    Room Lifecycle

    MethodDescription
    create_room(room_id?, metadata?, orchestration?)Create a room
    get_room(room_id)Get room by ID
    close_room(room_id)Close a room
    update_room_metadata(room_id, metadata)Update room metadata
    check_room_timers(room_id)Check timer transitions for one room
    check_all_timers()Check all room timers

    Channel Operations

    MethodDescription
    register_channel(channel)Register a channel
    attach_channel(room_id, channel_id, category?, access?, ...)Attach channel to room
    detach_channel(room_id, channel_id)Detach channel from room
    mute(room_id, channel_id)Mute a channel
    unmute(room_id, channel_id)Unmute a channel
    set_access(room_id, channel_id, access)Set channel access level

    Voice/Video

    MethodDescription
    join(room_id, channel_id, participant_id?, ...)Join voice/video session
    leave(session)Leave voice/video session
    transcribe(audio)Speech-to-text
    synthesize(text, voice?)Text-to-speech

    Inbound Pipeline

    MethodDescription
    process_inbound(message, room_id?)Process an inbound message

    Hooks

    MethodDescription
    hook(trigger, execution?, priority?, ...)Decorator to register a hook
    on(event_type)Decorator for framework events
    identity_hook(trigger, ...)Decorator for identity hooks
    on_delivery_status(fn)Decorator for delivery status
    add_room_hook(room_id, trigger, execution, fn, ...)Add room-scoped hook
    remove_room_hook(room_id, name)Remove room-scoped hook

    Realtime

    MethodDescription
    publish_typing(room_id, user_id, is_typing?)Typing indicator
    publish_presence(room_id, user_id, status)Presence update
    publish_reaction(room_id, user_id, target_event_id, emoji)Reaction
    publish_read_receipt(room_id, user_id, event_id)Read receipt
    subscribe_room(room_id, callback)Subscribe to ephemeral events
    unsubscribe_room(subscription_id)Unsubscribe

    Sources

    MethodDescription
    attach_source(channel_id, source, auto_restart?, ...)Attach event source
    detach_source(channel_id)Detach event source
    source_health(channel_id)Get source health

    Other

    MethodDescription
    delegate(room_id, agent_id, task, ...)Delegate to background agent
    send_greeting(room_id, channel_id?, greeting?, ...)Send greeting
    send_event(room_id, channel_id, content, ...)Send event directly
    get_timeline(room_id, offset?, limit?)Query event timeline
    close()Shutdown framework

    Provider Subpackage Imports

    AI Providers

    from roomkit.providers.anthropic.ai import AnthropicAIProvider
    from roomkit.providers.anthropic.config import AnthropicConfig
    from roomkit.providers.openai.ai import OpenAIAIProvider
    from roomkit.providers.openai.config import OpenAIConfig
    from roomkit.providers.gemini.ai import GeminiAIProvider
    from roomkit.providers.gemini.config import GeminiConfig
    from roomkit.providers.mistral.ai import MistralAIProvider
    from roomkit.providers.mistral.config import MistralConfig
    from roomkit.providers.ai.mock import MockAIProvider
    from roomkit.providers.ai.base import AIProvider, AIContext, AIResponse, AITool, AIToolCall
    

    SMS Providers

    from roomkit.providers.twilio.sms import TwilioSMSProvider
    from roomkit.providers.twilio.config import TwilioConfig
    from roomkit.providers.telnyx.sms import TelnyxSMSProvider
    from roomkit.providers.telnyx.config import TelnyxConfig
    from roomkit.providers.sinch.sms import SinchSMSProvider
    from roomkit.providers.sinch.config import SinchConfig
    from roomkit.providers.sms.mock import MockSMSProvider
    

    Voice (Lazy Loaders)

    from roomkit.voice import (
        get_deepgram_provider, get_deepgram_config,
        get_elevenlabs_provider, get_elevenlabs_config,
        get_sherpa_onnx_stt_provider, get_sherpa_onnx_tts_provider,
        get_local_audio_backend,
        get_fastrtc_backend,
        get_rtp_backend,
        get_sip_backend,
        get_gemini_live_provider,
        get_openai_realtime_provider,
        get_xai_realtime_provider,
        get_websocket_realtime_transport,
        get_speex_aec_provider,
        get_rnnoise_denoiser_provider,
    )
    

    Voice Mocks

    from roomkit.voice.backends.mock import MockVoiceBackend
    from roomkit.voice.stt.mock import MockSTTProvider
    from roomkit.voice.tts.mock import MockTTSProvider
    from roomkit.voice.realtime.mock import MockRealtimeProvider, MockRealtimeTransport
    

    Pipeline

    from roomkit.voice.pipeline import (
        AudioPipelineConfig, VADConfig,
        MockVADProvider, VADEvent, VADEventType,
        MockDenoiserProvider, MockDiarizationProvider,
        MockAGCProvider, MockAECProvider, MockDTMFDetector,
        MockAudioRecorder, MockTurnDetector, MockBackchannelDetector,
    )
    from roomkit.voice.interruption import InterruptionConfig, InterruptionStrategy
    from roomkit.voice.audio_frame import AudioFrame
    

    Orchestration

    from roomkit.orchestration.state import get_conversation_state, ConversationState
    from roomkit.orchestration.router import ConversationRouter, RoutingRule
    from roomkit.orchestration.pipeline import ConversationPipeline, PipelineStage
    from roomkit.orchestration.handoff import HandoffHandler, HandoffMemoryProvider
    

    Storage

    from roomkit.store.base import ConversationStore
    from roomkit.store.memory import InMemoryStore
    from roomkit.store.postgres import PostgresStore
    

    Content Types

    from roomkit.models.event import (
        TextContent, RichContent, MediaContent, AudioContent, VideoContent,
        LocationContent, CompositeContent, TemplateContent, SystemContent,
        EditContent, DeleteContent,
    )
    

    Premium Workflows

    Ready-to-use automation templates from the marketplace

    Browse marketplace →
    Multi-Tool Personal Assistant with Telegram, Grok-4, Gmail, Calendar & Memory
    n8n·$14.99· 66·★ 4.6
    Calculate Embodied Carbon (CO2) for Revit/IFC Models Using AI Classification
    n8n·$24.99· 15·★ 4.9
    Automate Customer Support with Grok-4, Google Docs, and Telegram
    n8n·$9.99· 88·★ 4.4
    Automate Twitter Content Analysis and Insights Delivery via Telegram
    n8n·$9.99· 42·★ 4.7
    X-Ray Pro: Twitter/X Insights with Grok AI & Telegram
    n8n·$22.99
    AI Reservation Reminder Calls via Twilio & Grok for Restaurants
    n8n·$22.99

    Stay up to date

    Get the latest Grok prompts, rules, and resources delivered to your inbox weekly.

    Neura Market LogoNeura Market

    Discover the best AI prompts, plugins, and resources for Grok and more.

    Content Types

    • Rules
    • Prompts
    • MCPs
    • Agents
    • Guides

    Platforms

    • ChatGPT Directory
    • Claude Directory
    • Gemini Directory
    • Cursor Directory
    • Grok Directory
    • Perplexity Directory
    • DeepSeek Directory
    • CoPilot Directory
    • Stable Diffusion Directory
    • Midjourney Directory
    • All Directories

    Resources

    • Blog
    • Documentation
    • Help Center
    • Marketplace

    Legal

    • Privacy Policy
    • Terms of Service

    © 2026 Neura Market. All rights reserved.

    |

    Not affiliated with any AI platform vendors.