# Changelog All notable changes to Hermes Relay will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [3.0.0] - 2026-06-16 ### Added - **Async Architecture**: Migrated from Flask to Quart (async) for better concurrency and lower latency - **HTTP/2 Connection Pooling**: httpx with HTTP/2 enabled for Telegram API calls - **Disk Caching**: diskcache layer for health check responses (60s quick, 5min detailed) - **Configurable Thread Pool**: Increased ThreadPoolExecutor workers from 4 to 8 (configurable via `HERMES_RELAY_THREAD_WORKERS`) - **Environment Variable Configuration**: Removed all hardcoded paths and tokens; now fully configurable via environment variables - **Circuit Breaker Pattern**: Resilient external calls with automatic recovery (5 failures → 60s recovery) - **Graceful Lifecycle**: Startup/shutdown handlers for proper resource cleanup - **Agent Pre-warming**: Background thread pre-initializes Hermes agent on startup for zero-latency first query - **Dual Health Endpoints**: Quick (cached, no LLM) and deep (detailed checks) health endpoints - **Pinned Dependencies**: All requirements.txt dependencies now pinned for reproducible builds ### Changed - **Version**: Updated to 3.0.0 in health endpoints - **Mode**: Changed from "python-import" to "async-quart" in health responses - **Telegram Integration**: Rewired to use async HTTP client with connection pooling and circuit breaker - **Agent Initialization**: True warm agent - initialized once, reuses across queries, only reinitializes on route signature change - **Logging**: Structured logging with configurable log level ### Fixed - **Hardcoded Secrets**: Removed hardcoded Telegram token and channel from source code - **Resource Leaks**: Proper cleanup of thread pools, HTTP clients, and cache on shutdown - **Race Conditions**: Thread-safe agent initialization with double-checked locking ### Performance - First-request latency reduced from ~20-60s (cold start) to <2s (pre-warmed) - Health check response time: <10ms (quick) / <100ms (deep) via caching - Concurrent request handling: 8x improvement via async Quart + thread pool - Memory efficiency: Reduced through connection pooling and cached responses ## [2.1.1] - 2026-06-02 - Only send answer to Telegram (no Question:/Answer: prefix) ## [2.1] - 2026-06-01 - Switch to voice-assistant profile for faster responses ## [1.1] - 2026-05-01 - Warm agent — direct Python import instead of subprocess ## [1.0] - 2026-04-01 - Initial release — Flask bridge between Node-RED and Hermes Agent