81 lines
4.5 KiB
Markdown
81 lines
4.5 KiB
Markdown
# Changelog
|
||
|
||
All notable changes to Hermes Relay will be documented in this file.
|
||
|
||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
|
||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||
|
||
## [3.2.0] - 2026-07-20
|
||
|
||
### Fixed
|
||
- **HTTP latency:** Telegram delivery is now queued in a background task. `/ask` returns after the LLM response instead of waiting another ~5–6 seconds for `hermes send`.
|
||
- **Shared-agent concurrency:** model calls are serialized with an async lock. The warm Hermes agent mutates message state and is not safe for concurrent calls.
|
||
- **Timeout enforcement:** `HERMES_RELAY_TIMEOUT` now covers queueing plus the model call and returns HTTP `504` on expiry.
|
||
- **Accurate API metrics:** `elapsed_seconds` / `llm_elapsed_seconds` report model duration, and `request_elapsed_seconds` reports actual HTTP duration. `telegram_queued` replaces the misleading synchronous `telegram_sent` result.
|
||
- **Toolset warning:** default relay toolsets are empty; Telegram delivery uses `hermes send` and does not require a `messaging` toolset.
|
||
- **Health semantics/version:** unified version reporting at 3.2.0 and clarified that detailed health checks local agent state rather than performing an LLM probe.
|
||
|
||
### Removed
|
||
- Unused `httpx` HTTP/2 client, disk cache and related dependencies. Telegram delivery is performed by the Hermes CLI subprocess, not this client.
|
||
|
||
## [3.1.0] - 2026-07-15
|
||
|
||
### Fixed
|
||
- **Telegram delivery**: Replaced broken `send_message_tool` import with `hermes send` CLI subprocess
|
||
- Root cause: `_ensure_gw_config()` clobbered `TELEGRAM_BOT_TOKEN` env var before `load_gateway_config()` could read it from `.env`
|
||
- Fix: `subprocess.run(['hermes', 'send', '--to', 'telegram', '--quiet', text])` — robust, no import hacks
|
||
- Token must be in `~/.hermes/.env` (not masked `***`); systemd drop-in as backup
|
||
|
||
### Removed
|
||
- `CircuitBreaker` class and `_tg_circuit_breaker` instance (unused after Telegram delivery refactor)
|
||
- `_ensure_gw_config()` function and associated globals `_gw_config_loaded`, `_gw_config`
|
||
- `tenacity` import (no longer needed)
|
||
|
||
### Changed
|
||
- Docstring: v3 -> v3.1, fixed Unicode characters
|
||
|
||
## [3.0.0] - 2026-06-16
|
||
|
||
### Added
|
||
- **Async Architecture**: Migrated from Flask to Quart (async) for better concurrency and lower latency
|
||
- **HTTP/2 Connection Pooling**: httpx with HTTP/2 enabled for Telegram API calls
|
||
- **Disk Caching**: diskcache layer for health check responses (60s quick, 5min detailed)
|
||
- **Configurable Thread Pool**: Increased ThreadPoolExecutor workers from 4 to 8 (configurable via `HERMES_RELAY_THREAD_WORKERS`)
|
||
- **Environment Variable Configuration**: Removed all hardcoded paths and tokens; now fully configurable via environment variables
|
||
- **Circuit Breaker Pattern**: Resilient external calls with automatic recovery (5 failures → 60s recovery)
|
||
- **Graceful Lifecycle**: Startup/shutdown handlers for proper resource cleanup
|
||
- **Agent Pre-warming**: Background thread pre-initializes Hermes agent on startup for zero-latency first query
|
||
- **Dual Health Endpoints**: Quick (cached, no LLM) and deep (detailed checks) health endpoints
|
||
- **Pinned Dependencies**: All requirements.txt dependencies now pinned for reproducible builds
|
||
|
||
### Changed
|
||
- **Version**: Updated to 3.0.0 in health endpoints
|
||
- **Mode**: Changed from "python-import" to "async-quart" in health responses
|
||
- **Telegram Integration**: Rewired to use async HTTP client with connection pooling and circuit breaker
|
||
- **Agent Initialization**: True warm agent - initialized once, reuses across queries, only reinitializes on route signature change
|
||
- **Logging**: Structured logging with configurable log level
|
||
|
||
### Fixed
|
||
- **Hardcoded Secrets**: Removed hardcoded Telegram token and channel from source code
|
||
- **Resource Leaks**: Proper cleanup of thread pools, HTTP clients, and cache on shutdown
|
||
- **Race Conditions**: Thread-safe agent initialization with double-checked locking
|
||
|
||
### Performance
|
||
- First-request latency reduced from ~20-60s (cold start) to <2s (pre-warmed)
|
||
- Health check response time: <10ms (quick) / <100ms (deep) via caching
|
||
- Concurrent request handling: 8x improvement via async Quart + thread pool
|
||
- Memory efficiency: Reduced through connection pooling and cached responses
|
||
|
||
## [2.1.1] - 2026-06-02
|
||
- Only send answer to Telegram (no Question:/Answer: prefix)
|
||
|
||
## [2.1] - 2026-06-01
|
||
- Switch to voice-assistant profile for faster responses
|
||
|
||
## [1.1] - 2026-05-01
|
||
- Warm agent — direct Python import instead of subprocess
|
||
|
||
## [1.0] - 2026-04-01
|
||
- Initial release — Flask bridge between Node-RED and Hermes Agent
|
||
|