fix: make relay fast and concurrency-safe

This commit is contained in:
2026-07-20 06:17:31 +02:00
parent 4eaa21f20c
commit 67160b2de0
5 changed files with 221 additions and 252 deletions
+13
View File
@@ -5,6 +5,19 @@ All notable changes to Hermes Relay will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [3.2.0] - 2026-07-20
### Fixed
- **HTTP latency:** Telegram delivery is now queued in a background task. `/ask` returns after the LLM response instead of waiting another ~5–6 seconds for `hermes send`.
- **Shared-agent concurrency:** model calls are serialized with an async lock. The warm Hermes agent mutates message state and is not safe for concurrent calls.
- **Timeout enforcement:** `HERMES_RELAY_TIMEOUT` now covers queueing plus the model call and returns HTTP `504` on expiry.
- **Accurate API metrics:** `elapsed_seconds` / `llm_elapsed_seconds` report model duration, and `request_elapsed_seconds` reports actual HTTP duration. `telegram_queued` replaces the misleading synchronous `telegram_sent` result.
- **Toolset warning:** default relay toolsets are empty; Telegram delivery uses `hermes send` and does not require a `messaging` toolset.
- **Health semantics/version:** unified version reporting at 3.2.0 and clarified that detailed health checks local agent state rather than performing an LLM probe.
### Removed
- Unused `httpx` HTTP/2 client, disk cache and related dependencies. Telegram delivery is performed by the Hermes CLI subprocess, not this client.
## [3.1.0] - 2026-07-15
### Fixed