When Logs Speak: Building an Agentic RCA System with Self-Evaluation Using Gemini

An incident-analysis system that detects log anomalies, generates root-cause analyses with Gemini, self-evaluates them, and notifies Slack.

Human and robot collaboratively examining log data in a server room

Representation Image

“What if logs didn’t just collect errors—but understood them, explained them, and improved their explanations on their own?”

👋 The Why

I’ve always been drawn to systems that can explain themselves—not just surface metrics. While today’s observability stacks offer visibility, there’s a gap when it comes to understanding incidents in real time.

Most observability setups stop at visibility.

You see metrics. You catch alerts. But when something breaks, it still takes human time, energy, and clarity of thought to:

  • Detect anomalies in behavior
  • Write a clean RCA
  • Review and publish it
  • Notify relevant stakeholders

It’s repeatable. It’s valuable—it teaches us a lot. And yet, it’s largely manual.

So I built something to change that: a system that could observe, reason, explain, and evaluate autonomously.

🧠 What I Built

A fully Dockerized Agentic RCA System that:

View the repository.

  1. Watches app behavior via logs
  2. Detects anomalies using MAD (Median Absolute Deviation)
  3. Writes human-readable RCAs via OpenAI
  4. Evaluates the quality using Gemini Flash (it was free at the time, and I was exploring it)
  5. Archives RCAs for auditability
  6. Notifies via Slack Webhooks

This isn’t just a pipeline. It’s a closed-loop agentic architecture—where each step is independent but purposeful.

📐 Design

Agentic RCA system design from test application and Fluent Bit through OpenSearch, extraction, LLM evaluation, and Slack notification

Basic representation of the idea

⚙️ Core Components

agentic_monitoring/
├── app/ # Simulated API app with realistic endpoints
├── fluentbit/ # Scrapes logs and forwards to OpenSearch
├── opensearch/ # Stores all logs, searchable via API
├── extractor/ # Detects anomalies using MAD
├── agent/ # Crafts RCA and refines it via Gemini
├── notifier/ # Sends RCA summary to Slack
├── docker-compose.yml

Each component is containerized for clean separation and easy testing.

🌐 API Simulation Layer

We simulate traffic with a realistic API surface:

  • /
  • /data
  • /metrics
  • /events
  • /notifications
  • /reports

Each endpoint emits logs with varying latency, status codes, and payload sizes—just like a production workload under moderate stress.

A script load_generator.py simulates the load on the machine with varying configurations.

🔁 End-to-End Flow

1. 📥 Log Ingestion

Logs are emitted from the test app. Fluent Bit tails them and ships to OpenSearch, where they’re indexed and timestamped.

2. 📊 Anomaly Detection via MAD

The extractor layer pulls logs from OpenSearch periodically and groups them by:

  • service
  • status_code
  • time window

We apply MAD (Median Absolute Deviation) to detect statistical anomalies in behavior. Once anomalies are found:

  • They’re written to a file called deltas.json
  • Invalid or unknown API hits are recorded too

JSON anomaly output identifying an unexpected reports endpoint and NotFound errors in the gateway service

System detecting invalid hits

3. 🧠 RCA Generation (with Evaluation Loop)

The agent layer picks up deltas.json and:

  • Sends it to OpenAI with a structured RCA prompt
  • Gets a response back (version 1 of the RCA)

Approved RCA JSON linking an unexpected reports endpoint to gateway errors and recommending an OpenAPI schema update

Approved RCA

Then comes the Gemini loop:

  • Gemini evaluates the RCA’s clarity, completeness, and tone
  • If it’s not good enough, we retry with Gemini’s feedback (up to 5 times)

Once Gemini is satisfied or retries are exhausted, the final output is saved to rca.json.

Gemini evaluation explaining why the generated RCA is clear, reasonable, actionable, and aligned with anomaly data

RCA evaluation by Gemini

4. 🗃️ Archiving & Versioning

At this stage:

  • Previous RCA files are archived in timestamped .json bundles
  • This ensures auditable, reproducible RCAs for every anomaly window

5. 🔔 Slack Notification

The notifier picks up rca.json and pushes a clear summary to Slack using a preconfigured webhook.

💡 Why MAD?

I tried other baselines like percentiles and z-scores. But MAD worked best in our case for:

  • Small sample sizes
  • Skewed distributions
  • Sudden latency or error spikes

This made it resilient to noise and ideal for short rolling windows of log data. You can use the other two methods by adjusting the environment variables, or contribute other methods to the codebase.

🧠 Why Gemini?

Adding Gemini made a huge difference. OpenAI did well generating RCAs—but Gemini Flash acted like a peer reviewer:

  • It flagged vague phrasing
  • Requested downstream impact details
  • Suggested actionable improvements

The result? RCAs that weren’t just machine-generated—but genuinely usable in postmortems and Slack threads.

🧪 Sample RCA Flow

{
 "timestamp": "2025-07-29T17:35:01Z",
 "service": "events",
 "original_rca": "High 5xx errors detected between 17:30–17:35 IST on /events endpoint due to a spike in malformed event payloads.",
 "gemini_feedback": "Clarify user impact. Include mitigation step.",
 "improved_rca": "Between 17:30–17:35 IST, the /events endpoint experienced a spike in 5xx errors due to malformed event payloads, impacting ~18% of users submitting real-time actions. The issue was mitigated by enforcing stricter schema validation and rate-limiting malformed POSTs."
}

Sample messages transmitted for demo.

Slack alert from the agentic AI bot reporting a possible systemic incident across multiple services

Bot reporting the errors when triggered on demand

Slack RCA alert identifying the unexpected reports endpoint as the cause of gateway and dependent-service errors

Detecting the issue at logs due to faulty endpoint usage

🧰 Technologies Used

+-------------------+----------------------------+
| Layer | Stack / Tool |
+-------------------+----------------------------+
| Log Collection | Fluent Bit |
| Storage & Search | OpenSearch |
| Anomaly Detection | Python + MAD logic |
| RCA Generation | OpenAI (gpt-4) |
| Evaluation | Gemini Flash API |
| Notification | Slack Webhook |
| Containerization | Docker + Compose |
| Archival Format | JSON (timestamped bundles) |
+-------------------+----------------------------+

🚀 What’s Next

I’m exploring:

  • Automated triggers for the extractor, agent, and notifier using cron jobs or an equivalent mechanism for real-time scenarios
  • Rule system under agent and notifier layer
  • Slack command to query past RCAs
  • Retrieval from archived RCAs using semantic search
  • Adding LangGraph for retry control and fallback reasoning
  • Building a Web UI for RCA playback and feedback scoring

🔍 Why This Matters

This isn’t just another automation tool.

It’s a system that:

  • Observes its environment
  • Reasons about failure
  • Learns from feedback
  • Communicates clearly
  • Documents responsibly

This architecture made us faster, sharper, and more auditable. It saved engineering hours and gave us confidence in the consistency of our RCA process.

And it’s just the beginning.

🐛 System Anomalies and Known Gaps

There are several issues in the code that I plan to improve over time.

  • I primarily used Docker Desktop for testing and development. This might be a limitation for some readers.
  • The extractor, agent, and notifier currently require manual triggers. I plan to automate them.
  • RCA archival function is not working as expected. To be fixed.
  • The code largely lacks docstrings and comments in some sections. I will improve this over time.
  • The docker-compose.yml file has intertwined volume mounts. I will fix this in the next version.
  • requirements.txt does not pin versions. I will fix this too.

🧘 Final Words

I built this to reduce toil—but ended up building a system that feels more like a teammate than a tool.

If you’re working in platform engineering, incident response, or observability—and you’re curious about building systems that think with you—this is your sign to start.

Logs don’t just have to scream. They can also explain.

Let’s Connect

If you’re exploring AI + observability or building agentic workflows, reach out. I’d love to chat, pair, or collaborate.

References:

  • Code at GitHub—link
  • Median Absolute Deviation—link
  • Gemini 2.0 flash details here
  • Digital art generated using ComfyUI. Read about it here

Originally published on Medium on July 28, 2025.