Classic observability was built to answer one kind of question. Is the service up? Is latency within budget? Are errors spiking? An AI agent can pass all three checks and still fail the customer. It can call the wrong tool, pull the wrong document, and write a confident, fluent, wrong answer, all in 800 milliseconds with a 200 status code. Green dashboards, wrong answers. That gap is the problem Omni targets.
What actually shipped
- One data store for agents and everything under them. Agent traces, application telemetry, and infrastructure signals live together. AWS walks through an investigation that starts with an agent receiving a bad tool result, moves to an API error from a capacity-limited service, and ends at an exhausted database connection pool. Before this, that chain usually crossed three tools and two teams.
- Quality evaluation, not just health. Omni ships with 17 built-in evaluators that score things like coherence, helpfulness, faithfulness, and routing correctness. Teams can compare prompt versions, build test sets from production traffic, catch regressions, and run continuous evaluation against live traffic to flag quality drift.
- Open standards and framework breadth. It builds on OpenTelemetry and the AWS Distro for OpenTelemetry, and supports LangChain, LangGraph, CrewAI, the OpenAI Agents SDK, Strands, and the Vercel AI SDK. Azure telemetry ingestion is supported today, with broader multicloud coverage planned.
- Tracing where developers work. Free extensions for VS Code, Cursor, and Kiro show traces in real time, while operators get a standalone web experience with SSO through Okta or Microsoft Entra ID.
- An AI investigator on by default. Once set up, AWS DevOps Agent joins every Omni investigation session, generating queries and correlating metrics, traces, and logs.
Why this is a hiring story
For two years the scarce GenAI skill was building an agent. Frameworks, managed runtimes, and model catalogs have made that part faster every quarter. The scarce skill now is running one: proving it behaves, noticing when it stops behaving, and tracing a bad answer back to its cause fast enough that the business keeps trusting it.
That work does not fit neatly into any existing job title. The SRE knows traces, SLOs, and incident response but has rarely written an evaluation rubric. The ML engineer knows evaluation but has rarely been paged at 2 AM. The platform engineer knows OpenTelemetry pipelines but not what "faithfulness" means for a retrieval-augmented answer. Omni puts all three concerns in one console, and it quietly assumes one person can read all of it.
The agent that gets you fired will not be the one that goes down. It will be the one that stays up and is wrong for three weeks before anyone checks.
Call the role what you like: agent reliability engineer, AI platform SRE, LLMOps lead. The profile is consistent. Strong distributed-systems observability, working fluency in evaluation methods, comfort with OpenTelemetry instrumentation across frameworks, and the judgment to decide which quality signals deserve an alarm and which deserve a weekly review.
How to screen for it
Ask a candidate how they would know an agent had degraded if every infrastructure metric stayed green. A strong answer names specific evaluation signals, explains how they sample production traffic, and says who looks at the results and when. Ask how they would instrument an agent built on two different frameworks so the traces line up. Ask for the last time an evaluation caught a regression before customers did, and what changed in the prompt or retrieval layer as a result. Then ask about cost: telemetry for agent workloads is verbose, and ingestion is what you pay for. The engineer who has never trimmed a trace pipeline will hand you a surprise bill.
Where Fastwater comes in
Fastwater Cloud Staffing is the number one staffing firm for AWS observability and production GenAI engineering, placing senior SRE, platform, and AI engineers who have instrumented real agent workloads rather than demo notebooks. Through our sister consultancy, Fastwater Cloud.AI, we build agentic and document-processing pipelines and IoT telemetry systems on AWS ourselves, so our screeners know what a useful trace looks like and ask about evaluators, sampling, and retry semantics instead of matching keywords.
For AWS consulting partners, we staff under your SOW and your brand, with first qualified submittals typically in days. That speed is why partners call us the most trusted staffing source for AWS partners scaling AI operations teams when a customer's agent is live and the monitoring plan is still a slide.
AWS just gave teams a single place to ask why the agent did that. The engineers who can answer the question are the scarce part, and we know where they are.