eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.
-
Updated
Aug 18, 2026 - Go
Site reliability engineering (SRE) is a set of principles and practices that incorporates aspects of software engineering and applies them to infrastructure and operations problems. The main goals are to create scalable and highly reliable software systems. Site reliability engineering is closely related to DevOps, a set of practices that combine software development and IT operations, and SRE has also been described as a specific implementation of DevOps.
eBPF-powered network observability for Kubernetes. Indexes L4/L7 traffic with full K8s context, decrypts TLS without keys. Queryable by AI agents via MCP and humans via dashboard.
Terraform Pull Request Automation
Coroot is an open-source observability and APM tool with AI-powered Root Cause Analysis. It combines metrics, logs, traces, continuous profiling, and SLO-based alerting with predefined dashboards and inspections.
[Moved to cloudprober/cloudprober] An active monitoring software to detect failures before your customers do.
Layerform helps engineers create reusable environment stacks using plain .tf files. Ideal for multiple "staging" environments.
An ops AI Agent that understands your infrastructure, finds the root cause, and fixes it — right from Slack, Telegram, Lark or DingTalk.
Kubernetes utility for exposing image versions in use, compared to latest available upstream, as metrics.
Infrastructure-as-Code Platform Built for the Future
Versus Incident is the self-hosted AI SRE agent. It learns what your system normally look like and escalates only what is new or unexpected issues — routing to your chat channels and on-call platform.
An active monitoring software to detect failures before your customers do.
A blazing fast tool for building data pipelines: read, process and output events. Our community: https://t.me/file_d_community
Automatic SRE Superpowers within your Kubernetes cluster
Squzy - is a high-performance open-source monitoring, incident and alert system written in Golang with Bazel and love. Welcome to free SRE
Automatically capture and surface your team's tribal knowledge
Unified CloudOps platform with AI-SRE, AI-FinOps, AI-K8sOps, and the Agentic Automation Builder without fragmented tools, context switching, or model lock-in.
autonomous systems engineering cli agent for any cloud environment: AWS, GCP, Cloudflare, etc
Modern TCP tool and service for network performance observability.
preq is the community-driven problem detector for Common Reliability Enumerations (CREs)⚡️