Files
hermes-relay/CHANGELOG.md
T

5.1 KiB
Raw Permalink Blame History

Changelog

All notable changes to Hermes Relay will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

[3.3.0] - 2026-08-16

Changed

  • HTTP-only response by default: /ask now returns the answer only to its HTTP caller. Relay responses are no longer delivered to Telegram automatically.
  • Explicit opt-in for Telegram: set HERMES_RELAY_TELEGRAM_DELIVERY_ENABLED=true only when asynchronous Telegram delivery is wanted. telegram_queued and /health now expose the active delivery state.
  • Hermes import compatibility: auto-detect the Hermes Git-checkout layout as well as the former virtualenv site-packages layout.

[3.2.0] - 2026-07-20

Fixed

  • HTTP latency: Telegram delivery is now queued in a background task. /ask returns after the LLM response instead of waiting another ~5–6 seconds for hermes send.
  • Shared-agent concurrency: model calls are serialized with an async lock. The warm Hermes agent mutates message state and is not safe for concurrent calls.
  • Timeout enforcement: HERMES_RELAY_TIMEOUT now covers queueing plus the model call and returns HTTP 504 on expiry.
  • Accurate API metrics: elapsed_seconds / llm_elapsed_seconds report model duration, and request_elapsed_seconds reports actual HTTP duration. telegram_queued replaces the misleading synchronous telegram_sent result.
  • Toolset warning: default relay toolsets are empty; Telegram delivery uses hermes send and does not require a messaging toolset.
  • Health semantics/version: unified version reporting at 3.2.0 and clarified that detailed health checks local agent state rather than performing an LLM probe.

Removed

  • Unused httpx HTTP/2 client, disk cache and related dependencies. Telegram delivery is performed by the Hermes CLI subprocess, not this client.

[3.1.0] - 2026-07-15

Fixed

  • Telegram delivery: Replaced broken send_message_tool import with hermes send CLI subprocess
    • Root cause: _ensure_gw_config() clobbered TELEGRAM_BOT_TOKEN env var before load_gateway_config() could read it from .env
    • Fix: subprocess.run(['hermes', 'send', '--to', 'telegram', '--quiet', text]) — robust, no import hacks
    • Token must be in ~/.hermes/.env (not masked ***); systemd drop-in as backup

Removed

  • CircuitBreaker class and _tg_circuit_breaker instance (unused after Telegram delivery refactor)
  • _ensure_gw_config() function and associated globals _gw_config_loaded, _gw_config
  • tenacity import (no longer needed)

Changed

  • Docstring: v3 -> v3.1, fixed Unicode characters

[3.0.0] - 2026-06-16

Added

  • Async Architecture: Migrated from Flask to Quart (async) for better concurrency and lower latency
  • HTTP/2 Connection Pooling: httpx with HTTP/2 enabled for Telegram API calls
  • Disk Caching: diskcache layer for health check responses (60s quick, 5min detailed)
  • Configurable Thread Pool: Increased ThreadPoolExecutor workers from 4 to 8 (configurable via HERMES_RELAY_THREAD_WORKERS)
  • Environment Variable Configuration: Removed all hardcoded paths and tokens; now fully configurable via environment variables
  • Circuit Breaker Pattern: Resilient external calls with automatic recovery (5 failures → 60s recovery)
  • Graceful Lifecycle: Startup/shutdown handlers for proper resource cleanup
  • Agent Pre-warming: Background thread pre-initializes Hermes agent on startup for zero-latency first query
  • Dual Health Endpoints: Quick (cached, no LLM) and deep (detailed checks) health endpoints
  • Pinned Dependencies: All requirements.txt dependencies now pinned for reproducible builds

Changed

  • Version: Updated to 3.0.0 in health endpoints
  • Mode: Changed from "python-import" to "async-quart" in health responses
  • Telegram Integration: Rewired to use async HTTP client with connection pooling and circuit breaker
  • Agent Initialization: True warm agent - initialized once, reuses across queries, only reinitializes on route signature change
  • Logging: Structured logging with configurable log level

Fixed

  • Hardcoded Secrets: Removed hardcoded Telegram token and channel from source code
  • Resource Leaks: Proper cleanup of thread pools, HTTP clients, and cache on shutdown
  • Race Conditions: Thread-safe agent initialization with double-checked locking

Performance

  • First-request latency reduced from ~20-60s (cold start) to <2s (pre-warmed)
  • Health check response time: <10ms (quick) / <100ms (deep) via caching
  • Concurrent request handling: 8x improvement via async Quart + thread pool
  • Memory efficiency: Reduced through connection pooling and cached responses

[2.1.1] - 2026-06-02

  • Only send answer to Telegram (no Question:/Answer: prefix)

[2.1] - 2026-06-01

  • Switch to voice-assistant profile for faster responses

[1.1] - 2026-05-01

  • Warm agent — direct Python import instead of subprocess

[1.0] - 2026-04-01

  • Initial release — Flask bridge between Node-RED and Hermes Agent