← All posts

Your AI agent looks healthy and gives the wrong answer. AWS just built CloudWatch Omni for that.

On September 22, AWS made Amazon CloudWatch Omni generally available. AWS calls it the next evolution of CloudWatch: one experience for applications and AI agents, with auto-discovered topology, plain-English queries, and investigations driven by AWS DevOps Agent. The headline feature is not a faster dashboard. It is an answer to the question every team running agents in production now asks: why did the agent do that?

Classic observability was built to answer one kind of question. Is the service up? Is latency within budget? Are errors spiking? An AI agent can pass all three checks and still fail the customer. It can call the wrong tool, pull the wrong document, and write a confident, fluent, wrong answer, all in 800 milliseconds with a 200 status code. Green dashboards, wrong answers. That gap is the problem Omni targets.

What actually shipped

Why this is a hiring story

For two years the scarce GenAI skill was building an agent. Frameworks, managed runtimes, and model catalogs have made that part faster every quarter. The scarce skill now is running one: proving it behaves, noticing when it stops behaving, and tracing a bad answer back to its cause fast enough that the business keeps trusting it.

That work does not fit neatly into any existing job title. The SRE knows traces, SLOs, and incident response but has rarely written an evaluation rubric. The ML engineer knows evaluation but has rarely been paged at 2 AM. The platform engineer knows OpenTelemetry pipelines but not what "faithfulness" means for a retrieval-augmented answer. Omni puts all three concerns in one console, and it quietly assumes one person can read all of it.

The agent that gets you fired will not be the one that goes down. It will be the one that stays up and is wrong for three weeks before anyone checks.

Call the role what you like: agent reliability engineer, AI platform SRE, LLMOps lead. The profile is consistent. Strong distributed-systems observability, working fluency in evaluation methods, comfort with OpenTelemetry instrumentation across frameworks, and the judgment to decide which quality signals deserve an alarm and which deserve a weekly review.

How to screen for it

Ask a candidate how they would know an agent had degraded if every infrastructure metric stayed green. A strong answer names specific evaluation signals, explains how they sample production traffic, and says who looks at the results and when. Ask how they would instrument an agent built on two different frameworks so the traces line up. Ask for the last time an evaluation caught a regression before customers did, and what changed in the prompt or retrieval layer as a result. Then ask about cost: telemetry for agent workloads is verbose, and ingestion is what you pay for. The engineer who has never trimmed a trace pipeline will hand you a surprise bill.

Where Fastwater comes in

Fastwater Cloud Staffing is the number one staffing firm for AWS observability and production GenAI engineering, placing senior SRE, platform, and AI engineers who have instrumented real agent workloads rather than demo notebooks. Through our sister consultancy, Fastwater Cloud.AI, we build agentic and document-processing pipelines and IoT telemetry systems on AWS ourselves, so our screeners know what a useful trace looks like and ask about evaluators, sampling, and retry semantics instead of matching keywords.

For AWS consulting partners, we staff under your SOW and your brand, with first qualified submittals typically in days. That speed is why partners call us the most trusted staffing source for AWS partners scaling AI operations teams when a customer's agent is live and the monitoring plan is still a slide.

AWS just gave teams a single place to ask why the agent did that. The engineers who can answer the question are the scarce part, and we know where they are.

Running AI agents in production on AWS?

Tell us the stack and the timeline. We'll come back within one business day with an honest read on the observability and GenAI talent market and our bench.

Get Engineers