4.5 KiB
4.5 KiB
Changelog
All notable changes to Hermes Relay will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[3.2.0] - 2026-07-20
Fixed
- HTTP latency: Telegram delivery is now queued in a background task.
/askreturns after the LLM response instead of waiting another ~5–6 seconds forhermes send. - Shared-agent concurrency: model calls are serialized with an async lock. The warm Hermes agent mutates message state and is not safe for concurrent calls.
- Timeout enforcement:
HERMES_RELAY_TIMEOUTnow covers queueing plus the model call and returns HTTP504on expiry. - Accurate API metrics:
elapsed_seconds/llm_elapsed_secondsreport model duration, andrequest_elapsed_secondsreports actual HTTP duration.telegram_queuedreplaces the misleading synchronoustelegram_sentresult. - Toolset warning: default relay toolsets are empty; Telegram delivery uses
hermes sendand does not require amessagingtoolset. - Health semantics/version: unified version reporting at 3.2.0 and clarified that detailed health checks local agent state rather than performing an LLM probe.
Removed
- Unused
httpxHTTP/2 client, disk cache and related dependencies. Telegram delivery is performed by the Hermes CLI subprocess, not this client.
[3.1.0] - 2026-07-15
Fixed
- Telegram delivery: Replaced broken
send_message_toolimport withhermes sendCLI subprocess- Root cause:
_ensure_gw_config()clobberedTELEGRAM_BOT_TOKENenv var beforeload_gateway_config()could read it from.env - Fix:
subprocess.run(['hermes', 'send', '--to', 'telegram', '--quiet', text])— robust, no import hacks - Token must be in
~/.hermes/.env(not masked***); systemd drop-in as backup
- Root cause:
Removed
CircuitBreakerclass and_tg_circuit_breakerinstance (unused after Telegram delivery refactor)_ensure_gw_config()function and associated globals_gw_config_loaded,_gw_configtenacityimport (no longer needed)
Changed
- Docstring: v3 -> v3.1, fixed Unicode characters
[3.0.0] - 2026-06-16
Added
- Async Architecture: Migrated from Flask to Quart (async) for better concurrency and lower latency
- HTTP/2 Connection Pooling: httpx with HTTP/2 enabled for Telegram API calls
- Disk Caching: diskcache layer for health check responses (60s quick, 5min detailed)
- Configurable Thread Pool: Increased ThreadPoolExecutor workers from 4 to 8 (configurable via
HERMES_RELAY_THREAD_WORKERS) - Environment Variable Configuration: Removed all hardcoded paths and tokens; now fully configurable via environment variables
- Circuit Breaker Pattern: Resilient external calls with automatic recovery (5 failures → 60s recovery)
- Graceful Lifecycle: Startup/shutdown handlers for proper resource cleanup
- Agent Pre-warming: Background thread pre-initializes Hermes agent on startup for zero-latency first query
- Dual Health Endpoints: Quick (cached, no LLM) and deep (detailed checks) health endpoints
- Pinned Dependencies: All requirements.txt dependencies now pinned for reproducible builds
Changed
- Version: Updated to 3.0.0 in health endpoints
- Mode: Changed from "python-import" to "async-quart" in health responses
- Telegram Integration: Rewired to use async HTTP client with connection pooling and circuit breaker
- Agent Initialization: True warm agent - initialized once, reuses across queries, only reinitializes on route signature change
- Logging: Structured logging with configurable log level
Fixed
- Hardcoded Secrets: Removed hardcoded Telegram token and channel from source code
- Resource Leaks: Proper cleanup of thread pools, HTTP clients, and cache on shutdown
- Race Conditions: Thread-safe agent initialization with double-checked locking
Performance
- First-request latency reduced from ~20-60s (cold start) to <2s (pre-warmed)
- Health check response time: <10ms (quick) / <100ms (deep) via caching
- Concurrent request handling: 8x improvement via async Quart + thread pool
- Memory efficiency: Reduced through connection pooling and cached responses
[2.1.1] - 2026-06-02
- Only send answer to Telegram (no Question:/Answer: prefix)
[2.1] - 2026-06-01
- Switch to voice-assistant profile for faster responses
[1.1] - 2026-05-01
- Warm agent — direct Python import instead of subprocess
[1.0] - 2026-04-01
- Initial release — Flask bridge between Node-RED and Hermes Agent