When Logs Speak: Building an Agentic RCA System with Self-Evaluation Using Gemini
An incident-analysis system that detects log anomalies, generates root-cause analyses with Gemini, self-evaluates them, and notifies Slack.

Representation Image
“What if logs didn’t just collect errors—but understood them, explained them, and improved their explanations on their own?”
👋 The Why
I’ve always been drawn to systems that can explain themselves—not just surface metrics. While today’s observability stacks offer visibility, there’s a gap when it comes to understanding incidents in real time.
Most observability setups stop at visibility.
You see metrics. You catch alerts. But when something breaks, it still takes human time, energy, and clarity of thought to:
- Detect anomalies in behavior
- Write a clean RCA
- Review and publish it
- Notify relevant stakeholders
It’s repeatable. It’s valuable—it teaches us a lot. And yet, it’s largely manual.
So I built something to change that: a system that could observe, reason, explain, and evaluate autonomously.
🧠 What I Built
A fully Dockerized Agentic RCA System that:
- Watches app behavior via logs
- Detects anomalies using MAD (Median Absolute Deviation)
- Writes human-readable RCAs via OpenAI
- Evaluates the quality using Gemini Flash (it was free at the time, and I was exploring it)
- Archives RCAs for auditability
- Notifies via Slack Webhooks
This isn’t just a pipeline. It’s a closed-loop agentic architecture—where each step is independent but purposeful.
📐 Design

Basic representation of the idea
⚙️ Core Components
agentic_monitoring/
├── app/ # Simulated API app with realistic endpoints
├── fluentbit/ # Scrapes logs and forwards to OpenSearch
├── opensearch/ # Stores all logs, searchable via API
├── extractor/ # Detects anomalies using MAD
├── agent/ # Crafts RCA and refines it via Gemini
├── notifier/ # Sends RCA summary to Slack
├── docker-compose.yml
Each component is containerized for clean separation and easy testing.
🌐 API Simulation Layer
We simulate traffic with a realistic API surface:
- /
- /data
- /metrics
- /events
- /notifications
- /reports
Each endpoint emits logs with varying latency, status codes, and payload sizes—just like a production workload under moderate stress.
A script load_generator.py simulates the load on the machine with varying configurations.
🔁 End-to-End Flow
1. 📥 Log Ingestion
Logs are emitted from the test app. Fluent Bit tails them and ships to OpenSearch, where they’re indexed and timestamped.
2. 📊 Anomaly Detection via MAD
The extractor layer pulls logs from OpenSearch periodically and groups them by:
- service
- status_code
- time window
We apply MAD (Median Absolute Deviation) to detect statistical anomalies in behavior. Once anomalies are found:
- They’re written to a file called deltas.json
- Invalid or unknown API hits are recorded too

System detecting invalid hits
3. 🧠 RCA Generation (with Evaluation Loop)
The agent layer picks up deltas.json and:
- Sends it to OpenAI with a structured RCA prompt
- Gets a response back (version 1 of the RCA)

Approved RCA
Then comes the Gemini loop:
- Gemini evaluates the RCA’s clarity, completeness, and tone
- If it’s not good enough, we retry with Gemini’s feedback (up to 5 times)
Once Gemini is satisfied or retries are exhausted, the final output is saved to rca.json.

RCA evaluation by Gemini
4. 🗃️ Archiving & Versioning
At this stage:
- Previous RCA files are archived in timestamped .json bundles
- This ensures auditable, reproducible RCAs for every anomaly window
5. 🔔 Slack Notification
The notifier picks up rca.json and pushes a clear summary to Slack using a preconfigured webhook.
💡 Why MAD?
I tried other baselines like percentiles and z-scores. But MAD worked best in our case for:
- Small sample sizes
- Skewed distributions
- Sudden latency or error spikes
This made it resilient to noise and ideal for short rolling windows of log data. You can use the other two methods by adjusting the environment variables, or contribute other methods to the codebase.
🧠 Why Gemini?
Adding Gemini made a huge difference. OpenAI did well generating RCAs—but Gemini Flash acted like a peer reviewer:
- It flagged vague phrasing
- Requested downstream impact details
- Suggested actionable improvements
The result? RCAs that weren’t just machine-generated—but genuinely usable in postmortems and Slack threads.
🧪 Sample RCA Flow
{
"timestamp": "2025-07-29T17:35:01Z",
"service": "events",
"original_rca": "High 5xx errors detected between 17:30–17:35 IST on /events endpoint due to a spike in malformed event payloads.",
"gemini_feedback": "Clarify user impact. Include mitigation step.",
"improved_rca": "Between 17:30–17:35 IST, the /events endpoint experienced a spike in 5xx errors due to malformed event payloads, impacting ~18% of users submitting real-time actions. The issue was mitigated by enforcing stricter schema validation and rate-limiting malformed POSTs."
}
Sample messages transmitted for demo.

Bot reporting the errors when triggered on demand

Detecting the issue at logs due to faulty endpoint usage
🧰 Technologies Used
+-------------------+----------------------------+
| Layer | Stack / Tool |
+-------------------+----------------------------+
| Log Collection | Fluent Bit |
| Storage & Search | OpenSearch |
| Anomaly Detection | Python + MAD logic |
| RCA Generation | OpenAI (gpt-4) |
| Evaluation | Gemini Flash API |
| Notification | Slack Webhook |
| Containerization | Docker + Compose |
| Archival Format | JSON (timestamped bundles) |
+-------------------+----------------------------+
🚀 What’s Next
I’m exploring:
- Automated triggers for the extractor, agent, and notifier using cron jobs or an equivalent mechanism for real-time scenarios
- Rule system under agent and notifier layer
- Slack command to query past RCAs
- Retrieval from archived RCAs using semantic search
- Adding LangGraph for retry control and fallback reasoning
- Building a Web UI for RCA playback and feedback scoring
🔍 Why This Matters
This isn’t just another automation tool.
It’s a system that:
- Observes its environment
- Reasons about failure
- Learns from feedback
- Communicates clearly
- Documents responsibly
This architecture made us faster, sharper, and more auditable. It saved engineering hours and gave us confidence in the consistency of our RCA process.
And it’s just the beginning.
🐛 System Anomalies and Known Gaps
There are several issues in the code that I plan to improve over time.
- I primarily used Docker Desktop for testing and development. This might be a limitation for some readers.
- The extractor, agent, and notifier currently require manual triggers. I plan to automate them.
- RCA archival function is not working as expected. To be fixed.
- The code largely lacks docstrings and comments in some sections. I will improve this over time.
- The
docker-compose.ymlfile has intertwined volume mounts. I will fix this in the next version. requirements.txtdoes not pin versions. I will fix this too.
🧘 Final Words
I built this to reduce toil—but ended up building a system that feels more like a teammate than a tool.
If you’re working in platform engineering, incident response, or observability—and you’re curious about building systems that think with you—this is your sign to start.
Logs don’t just have to scream. They can also explain.
Let’s Connect
If you’re exploring AI + observability or building agentic workflows, reach out. I’d love to chat, pair, or collaborate.
References:
- Code at GitHub—link
- Median Absolute Deviation—link
- Gemini 2.0 flash details here
- Digital art generated using ComfyUI. Read about it here
Originally published on Medium on July 28, 2025.