The harness you can audit and control.
OpenRemedy gives platform and SRE teams a documented, per-server trust gradient and an append-only audit trail β not a black-box autonomy switch.
Sustained CPU on api-prod-03
opened by daemon Β· 00:00 ago
Agent pipeline startingβ¦
Atlas is being assigned to the incident.
The cost of manual incident response
What a 3am page already costs you.
A senior engineer spends 20 to 40 minutes proving which known fix applies before acting. That cost exists today, in every on-call rotation, whether or not you use OpenRemedy.
Illustrative, not a customer result β OpenRemedy has no external customers yet, by design. This models the cost every SRE team already carries.
Engineer time per incident, before OpenRemedy
Median time OpenRemedy's pipeline takes for the same investigation
Use cases
Built for what's actually running in production.
Docker fleets in production
A dedicated Docker SRE specialist diagnoses container-host incidents with real tools β disk usage, container inspection, process-level detail β not a generic runbook.
Hypervisor & Proxmox operations
A separate domain specialist handles hypervisor-layer incidents distinctly from container-host ones β the right agent, the right tools, for the right kind of server.
Planned maintenance without surprise downtime
Maintenance plans run under rolling, canary, or ring strategies with per-step approval gates β the same operator-control model that governs incident remediation.
Compliance & audit
Every state-changing action writes an append-only audit row. The trust ladder records who granted autonomy, when, and based on what evidence.
Sustained CPU on api-prod-03
opened 41 seconds ago Β· daemon Β· resolved automatically
Triage Β· Atlas
Pattern matches a recurring spike from the daily ingest job. Routed to the diagnose stage.
Diagnose Β· Forge
Captured a fresh top -bn1 snapshot. Top consumer: python ingest_worker.py at 89% (expected behaviour during ingest).
Resolved Β· Default SRE
Transient β load returning to baseline. No action taken. Marked for monitoring.
Approval requested
Restart nginx on edge-04?
Proposed recipe
systemd-restart-service
risk: medium Β· trust: supervised Β· expected duration: ~8 s
Why this fix
nginx.service has been in failed for 4 min. Last 12 lines of the journal show recurring SIGSEGV after a config reload. A clean restart is the standard remedy and reproduces past resolutions on this server.
Deployment & security posture
Self-hosted control, gated execution.
The entire control plane β database, queue, secret store, model gateway β is built to run inside your own perimeter. Every remediation passes through the trust ladder and a two-stage approval gate before it touches production.
nginx restart on edge-04
Tue, 14:08 UTC Β· 12 minutes total Β· approved by alberto@β¦
- What happened
- nginx.service crashed with SIGSEGV after a config reload at 14:01. The daemon detected the failed state within 12 seconds and opened an incident.
- Root cause
- A reload pulled in a partially-written /etc/nginx/sites-enabled/api.conf. The deploy pipeline had no atomic-write step.
- What we did
- Approved the systemd-restart-service recipe at 14:09. Service back to active (running) in 6 s. Verified with three follow-up health probes.
- Follow-up
- Open ticket against the deploy pipeline to add atomic file-replace. Add a proactive policy to alert on partial config files.
Integrates with what you have
Not a new island to manage.
Plugins and webhooks connect OpenRemedy to Slack, Microsoft Teams, ServiceNow, Jira, and generic webhook endpoints β incidents and approvals show up where your team already works.
Talk to us about your fleet.
OpenRemedy is in private testing. Join the waitlist to be among the first enterprise operators we work with.