# Changelog All notable changes to Hermes Relay will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [3.2.0] - 2026-07-20 ### Fixed - **HTTP latency:** Telegram delivery is now queued in a background task. `/ask` returns after the LLM response instead of waiting another ~5–6 seconds for `hermes send`. - **Shared-agent concurrency:** model calls are serialized with an async lock. The warm Hermes agent mutates message state and is not safe for concurrent calls. - **Timeout enforcement:** `HERMES_RELAY_TIMEOUT` now covers queueing plus the model call and returns HTTP `504` on expiry. - **Accurate API metrics:** `elapsed_seconds` / `llm_elapsed_seconds` report model duration, and `request_elapsed_seconds` reports actual HTTP duration. `telegram_queued` replaces the misleading synchronous `telegram_sent` result. - **Toolset warning:** default relay toolsets are empty; Telegram delivery uses `hermes send` and does not require a `messaging` toolset. - **Health semantics/version:** unified version reporting at 3.2.0 and clarified that detailed health checks local agent state rather than performing an LLM probe. ### Removed - Unused `httpx` HTTP/2 client, disk cache and related dependencies. Telegram delivery is performed by the Hermes CLI subprocess, not this client. ## [3.1.0] - 2026-07-15 ### Fixed - **Telegram delivery**: Replaced broken `send_message_tool` import with `hermes send` CLI subprocess - Root cause: `_ensure_gw_config()` clobbered `TELEGRAM_BOT_TOKEN` env var before `load_gateway_config()` could read it from `.env` - Fix: `subprocess.run(['hermes', 'send', '--to', 'telegram', '--quiet', text])` — robust, no import hacks - Token must be in `~/.hermes/.env` (not masked `***`); systemd drop-in as backup ### Removed - `CircuitBreaker` class and `_tg_circuit_breaker` instance (unused after Telegram delivery refactor) - `_ensure_gw_config()` function and associated globals `_gw_config_loaded`, `_gw_config` - `tenacity` import (no longer needed) ### Changed - Docstring: v3 -> v3.1, fixed Unicode characters ## [3.0.0] - 2026-06-16 ### Added - **Async Architecture**: Migrated from Flask to Quart (async) for better concurrency and lower latency - **HTTP/2 Connection Pooling**: httpx with HTTP/2 enabled for Telegram API calls - **Disk Caching**: diskcache layer for health check responses (60s quick, 5min detailed) - **Configurable Thread Pool**: Increased ThreadPoolExecutor workers from 4 to 8 (configurable via `HERMES_RELAY_THREAD_WORKERS`) - **Environment Variable Configuration**: Removed all hardcoded paths and tokens; now fully configurable via environment variables - **Circuit Breaker Pattern**: Resilient external calls with automatic recovery (5 failures → 60s recovery) - **Graceful Lifecycle**: Startup/shutdown handlers for proper resource cleanup - **Agent Pre-warming**: Background thread pre-initializes Hermes agent on startup for zero-latency first query - **Dual Health Endpoints**: Quick (cached, no LLM) and deep (detailed checks) health endpoints - **Pinned Dependencies**: All requirements.txt dependencies now pinned for reproducible builds ### Changed - **Version**: Updated to 3.0.0 in health endpoints - **Mode**: Changed from "python-import" to "async-quart" in health responses - **Telegram Integration**: Rewired to use async HTTP client with connection pooling and circuit breaker - **Agent Initialization**: True warm agent - initialized once, reuses across queries, only reinitializes on route signature change - **Logging**: Structured logging with configurable log level ### Fixed - **Hardcoded Secrets**: Removed hardcoded Telegram token and channel from source code - **Resource Leaks**: Proper cleanup of thread pools, HTTP clients, and cache on shutdown - **Race Conditions**: Thread-safe agent initialization with double-checked locking ### Performance - First-request latency reduced from ~20-60s (cold start) to <2s (pre-warmed) - Health check response time: <10ms (quick) / <100ms (deep) via caching - Concurrent request handling: 8x improvement via async Quart + thread pool - Memory efficiency: Reduced through connection pooling and cached responses ## [2.1.1] - 2026-06-02 - Only send answer to Telegram (no Question:/Answer: prefix) ## [2.1] - 2026-06-01 - Switch to voice-assistant profile for faster responses ## [1.1] - 2026-05-01 - Warm agent — direct Python import instead of subprocess ## [1.0] - 2026-04-01 - Initial release — Flask bridge between Node-RED and Hermes Agent