Marketplace / Blueprint 13
LIVEOperations / observability

AI Operations Center

Bring metrics, logs, traces and platform health into one operating picture.

Review my build

Interactive platform map

Architecture in context

The focused blueprint, its required foundation, and declared recommendations.

Core Selected Automatic Required path
Blueprint 13 / Operations

AI Operations Center

Bring metrics, logs, traces and platform health into one operating picture.

3 vCPU7 GB RAM8 services
01 / Problem

What this replaces

Complex AI platforms fail across multiple layers, making root cause analysis difficult when data is scattered across containers and services.

02 / Outcome

What your team gains

A unified observability stack for infrastructure and applications with alerting, correlation and long-retention metrics.

03 / Capability

What is inside the blueprint

VictoriaMetrics Prometheus-compatible remote storage
Long retention
PromQL/MetricsQL-oriented monitoring workflows
vmalert alerting and recording rules
Grafana dashboards and Explore
Grafana multi-data-source queries and alert rules
Loki highly compressed label-indexed logging
LogQL search and metrics-from-logs
Alloy collection of metrics, logs, traces and profiles
Prometheus and OpenTelemetry pipelines
OTLP receive/export pipelines
Tempo distributed tracing
OTLP/Jaeger/Zipkin ingestion
TraceQL search
Trace-to-logs and trace-to-metrics correlation
Service graphs and RED metrics in Grafana
PNG panel rendering
Infrastructure status dashboard
04 / Architecture

How it fits the platform

Services/host -> Alloy/Prometheus/OTLP/log files -> VictoriaMetrics + Loki + Tempo -> Grafana -> vmalert/Grafana alerts -> Telegram/reports.

Included services

VictoriaMetrics, vmalert, Grafana, renderer, Loki, Alloy, Tempo, ops-dashboard

Platform requirements
05 / Delivery

From prerequisites to operation

Prerequisites
  1. Telemetry endpoints/log paths
  2. Storage sizing
  3. Dashboard and alert objectives
Deployment
  1. Deploy metrics/log/trace backends
  2. Configure Alloy pipelines and scrapes
  3. Register Grafana data sources
  4. Import/create dashboards
  5. Create alert rules and notifier path
Configuration
  1. Retention
  2. Scrape targets
  3. Log labels
  4. OTLP endpoints
  5. Trace sampling
  6. Dashboards
  7. Alert thresholds
Operations
  1. Storage/cardinality growth
  2. Alert delivery
  3. Collector health
  4. Trace cost/benefit
  5. Dashboard lifecycle
06 / Combinations

What this unlocks with other layers

AI Operations Center + MCP Agent Hub + Persistent AI Memory + Knowledge Graph & Institutional Intelligence

AI SRE Foundation

Agents can reason over the same telemetry and history humans use for diagnosis.

Threat Detection & Runtime Security + AI Operations Center

Security Operations Story

Edge attacks, software vulnerabilities and runtime behavior feed one alert and response narrative.

07 / Technology

Technology behind this capability

VictoriaMetricsRUNNING - v1.115.0 in censusvmalertRUNNINGGrafanaRUNNING - 13.1.0 in censusGrafana Image RendererRUNNINGLokiRUNNING - 14-day corpus policyGrafana AlloyRUNNING - v1.12.2 in censusTempoRUNNING - 2.9.0 in censusops-dashboardAICORTEX native