Files
hermes-relay/CHANGELOG.md
T

88 lines
5.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Changelog
All notable changes to Hermes Relay will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [3.3.0] - 2026-08-16
### Changed
- **HTTP-only response by default:** `/ask` now returns the answer only to its HTTP caller. Relay responses are no longer delivered to Telegram automatically.
- **Explicit opt-in for Telegram:** set `HERMES_RELAY_TELEGRAM_DELIVERY_ENABLED=true` only when asynchronous Telegram delivery is wanted. `telegram_queued` and `/health` now expose the active delivery state.
- **Hermes import compatibility:** auto-detect the Hermes Git-checkout layout as well as the former virtualenv site-packages layout.
## [3.2.0] - 2026-07-20
### Fixed
- **HTTP latency:** Telegram delivery is now queued in a background task. `/ask` returns after the LLM response instead of waiting another ~5–6 seconds for `hermes send`.
- **Shared-agent concurrency:** model calls are serialized with an async lock. The warm Hermes agent mutates message state and is not safe for concurrent calls.
- **Timeout enforcement:** `HERMES_RELAY_TIMEOUT` now covers queueing plus the model call and returns HTTP `504` on expiry.
- **Accurate API metrics:** `elapsed_seconds` / `llm_elapsed_seconds` report model duration, and `request_elapsed_seconds` reports actual HTTP duration. `telegram_queued` replaces the misleading synchronous `telegram_sent` result.
- **Toolset warning:** default relay toolsets are empty; Telegram delivery uses `hermes send` and does not require a `messaging` toolset.
- **Health semantics/version:** unified version reporting at 3.2.0 and clarified that detailed health checks local agent state rather than performing an LLM probe.
### Removed
- Unused `httpx` HTTP/2 client, disk cache and related dependencies. Telegram delivery is performed by the Hermes CLI subprocess, not this client.
## [3.1.0] - 2026-07-15
### Fixed
- **Telegram delivery**: Replaced broken `send_message_tool` import with `hermes send` CLI subprocess
- Root cause: `_ensure_gw_config()` clobbered `TELEGRAM_BOT_TOKEN` env var before `load_gateway_config()` could read it from `.env`
- Fix: `subprocess.run(['hermes', 'send', '--to', 'telegram', '--quiet', text])` — robust, no import hacks
- Token must be in `~/.hermes/.env` (not masked `***`); systemd drop-in as backup
### Removed
- `CircuitBreaker` class and `_tg_circuit_breaker` instance (unused after Telegram delivery refactor)
- `_ensure_gw_config()` function and associated globals `_gw_config_loaded`, `_gw_config`
- `tenacity` import (no longer needed)
### Changed
- Docstring: v3 -> v3.1, fixed Unicode characters
## [3.0.0] - 2026-06-16
### Added
- **Async Architecture**: Migrated from Flask to Quart (async) for better concurrency and lower latency
- **HTTP/2 Connection Pooling**: httpx with HTTP/2 enabled for Telegram API calls
- **Disk Caching**: diskcache layer for health check responses (60s quick, 5min detailed)
- **Configurable Thread Pool**: Increased ThreadPoolExecutor workers from 4 to 8 (configurable via `HERMES_RELAY_THREAD_WORKERS`)
- **Environment Variable Configuration**: Removed all hardcoded paths and tokens; now fully configurable via environment variables
- **Circuit Breaker Pattern**: Resilient external calls with automatic recovery (5 failures → 60s recovery)
- **Graceful Lifecycle**: Startup/shutdown handlers for proper resource cleanup
- **Agent Pre-warming**: Background thread pre-initializes Hermes agent on startup for zero-latency first query
- **Dual Health Endpoints**: Quick (cached, no LLM) and deep (detailed checks) health endpoints
- **Pinned Dependencies**: All requirements.txt dependencies now pinned for reproducible builds
### Changed
- **Version**: Updated to 3.0.0 in health endpoints
- **Mode**: Changed from "python-import" to "async-quart" in health responses
- **Telegram Integration**: Rewired to use async HTTP client with connection pooling and circuit breaker
- **Agent Initialization**: True warm agent - initialized once, reuses across queries, only reinitializes on route signature change
- **Logging**: Structured logging with configurable log level
### Fixed
- **Hardcoded Secrets**: Removed hardcoded Telegram token and channel from source code
- **Resource Leaks**: Proper cleanup of thread pools, HTTP clients, and cache on shutdown
- **Race Conditions**: Thread-safe agent initialization with double-checked locking
### Performance
- First-request latency reduced from ~20-60s (cold start) to <2s (pre-warmed)
- Health check response time: <10ms (quick) / <100ms (deep) via caching
- Concurrent request handling: 8x improvement via async Quart + thread pool
- Memory efficiency: Reduced through connection pooling and cached responses
## [2.1.1] - 2026-06-02
- Only send answer to Telegram (no Question:/Answer: prefix)
## [2.1] - 2026-06-01
- Switch to voice-assistant profile for faster responses
## [1.1] - 2026-05-01
- Warm agent — direct Python import instead of subprocess
## [1.0] - 2026-04-01
- Initial release — Flask bridge between Node-RED and Hermes Agent