Product · heronsentry

HeronSentry

HeronSentry is a standalone observability plane for AI Agents: capturing runtime events via standard OTLP to deliver call chain traces, cost aggregation, alerts, and performance exports without depending on other suite products.

HeronSentry (Observability Platform) architecture diagram for enterprise digital workforce infrastructure
HeronSentry logo

HeronSentry

Observability PlatformStandalone AI Agent observability plane: tracing, costs, alerts, and performance analysis

What Agents do, how well they perform, and what they cost—visible at a glance via standard OTLP.

OTLP
Open protocol; semantic fields mapped via profiles
Full Ingestion
Unsampled by default
Synchronous Ingest
Persisted within request lifecycle
Standalone
Deployable with zero suite dependencies

Typical pains

  1. Troubleshooting broken Agent workflows relies on raw logs without distributed call chains
  2. Costs remain invisible until month-end invoices arrive after budgets are blown
  3. No visibility into whether specific Agents are healthy or experiencing abuse

Core value

Runtime alerts with optional webhook fan-out

Error rate, cost spikes, ingestion backlog, and high-frequency tool calls can fan out via webhook once an endpoint is configured; latency spikes and orphan spans skip the webhook and stay on the alerts view.

Alerts: Multi-source runtime monitoring and trace inspection

Latency, error rate and truncation for each agent

HeronSentry rolls up response latency, streaming time to first token (TTFT), error rate and truncation rate for each agent, along with the model and provider it used most recently. AULO shows them in the agent detail panel of its console.

Console agent detail: model, response latency, time to first token, error rate and truncation

Tokens and spend by model, host and agent

The token usage and spend HeronSentry collects are grouped in AULO's resource overview by model, by host and by agent, so one screen shows which model, and which agent on which machine, uses the most.

Console resource overview: token usage and spend by model, host and agent

Project spend by model and agent

HeronSentry attributes a project's tokens and spend to each model and each agent, and splits tokens into input, cached and output. Budget utilization comes from PathPilot.

Project cost tab: token usage and spend by model and agent

Alert rule thresholds and switches

HeronSentry's alert rules are listed in AULO with their conditions, thresholds and severity. Filter them by source and switch each one on or off.

Alert rules: conditions, thresholds, severity and on/off switches

Capabilities

Ingestion

Standalone DeploymentZero suite dependencies; built-in ingestion, PostgreSQL storage, alert engine, and debug view
Standard OTLP IngestionReceives OTLP traces / metrics / logs (HTTP and gRPC); built-in profiles for OpenClaw, OpenLLMetry, Claude Code, Codex, Gemini CLI, and NodalOS; other languages can report Spans via HTTP Push
OTLP Ingest Rate LimitOTLP HTTP and gRPC can share a rate-limit bucket (per API key or peer; off by default); this does not cover native ingest

Does / does not

Does
  • Standard OTLP ingestion: traces, metrics, and logs persisted in real time
  • Distributed call-chain tracing: reconstruct Agent reasoning, tool calls, and model interaction topology
  • Granular token and cache cost accounting: split cache read/write for real-time cost attribution
  • Multi-dimensional alerting: error rate, cost spikes, backlog, and anomalous calls, with optional webhooks
  • Agent reliability metrics: TTFT, truncation rate, and tool-rejection analysis
  • Analytics and audit data export: metrics and audit feeds for OwlAudit and operators
Does not
  • Does not serve as a generic enterprise full-text log search cluster (focuses on Agent-runtime OTLP traces and structured logs)
  • Does not replace host/network infrastructure monitoring (focuses on Agent application and model-interaction semantics)
  • Does not execute active runtime interventions (read-only observability and alerts; blocking belongs to OwlAudit / NodalOS)

By role

Technology leaders

Standard OTLP ingestion avoids framework lock-in; coexists with existing observability stacks without claiming backend reuse

Compliance & risk

OwlAudit pulls operational telemetry exports from this service

Business leaders

Clear cost attribution and measurable ROI for enterprise Agent deployments

Integrations

NodalOS OTLP push (requires self-telemetry reporting enabled on the NodalOS side)AULO alerts viewPathPilot cost writeback and runtime health (error rate / P95 / timeouts)OwlAudit export ingestion for traces, costs, and alerts

Related products

Next steps

SYSTEM READYproduct/heronsentry