Observability
Query your environments' logs and metrics with SQL.
Every deployed environment streams logs and metrics into a queryable store. Use specific query to run SQL against this data when debugging production incidents, investigating regressions, or doing ad-hoc analytics.
Running a query
# Inline
specific query "SELECT count() FROM observability.logs"
# Against a specific environment
specific query --environment staging "SELECT * FROM observability.logs LIMIT 5"
# From a file via stdin
cat queries/p99.sql | specific query
# Machine-readable output for agents and scripts
specific query --format jsonl "SELECT Timestamp, Body FROM observability.logs LIMIT 5"
Flags:
| Flag | Description |
|---|---|
-e, --environment <name|id> |
Target environment by name or ID (defaults to the current one). Run specific status to list environments. |
--db <name> |
Run the query against one of the environment’s Postgres databases (read-only) instead of the observability store. The SQL is standard Postgres, and the schema is your application’s tables; see Postgres. |
--format <table|json|jsonl|csv> |
Output format. Defaults to table on a terminal, jsonl when piped. |
Queries are automatically scoped to the selected environment: you can’t see data from other environments and don’t need to filter on environment yourself.
Queries are read-only: INSERT, UPDATE, DELETE, and DDL are rejected. Each query is capped at 30 seconds of execution time, returns at most 100,000 rows, and the SQL string is limited to 50,000 characters.
The table format truncates wide values for readability. Agents and scripts should use jsonl, json, or csv for complete, untruncated values. Notices are printed to stderr so stdout remains parseable.
Choosing an environment
specific query runs against exactly one environment. Run specific status to discover every environment and its ID. The output has two sections:
- Environments - long-lived environments such as
productionandstaging. - Preview environments - ephemeral environments created for pull requests (plus any manually created previews), each with its name,
env_...ID, PR number and title, deployed URL, and expiry.
Preview environments stream logs and metrics exactly like long-lived ones. To inspect one, pass its name or ID to --environment; that’s the only way to reach a preview’s data, since queries are auto-scoped:
specific query --environment preview-a1b2c3d4 \
"SELECT Timestamp, ServiceName, Body
FROM observability.logs
WHERE SeverityNumber >= 17
ORDER BY Timestamp DESC
LIMIT 50"
Schema
Four views are exposed in the observability database: logs, metrics, traces and events. Column names follow OpenTelemetry’s ClickHouse exporter conventions and are case-sensitive, so use SeverityText, ServiceName, and ResourceAttributes exactly as shown. If unsure, run DESCRIBE observability.logs.
observability.logs
Unified log stream from services. Retention: up to 30 days (your plan may restrict how far back queries can read; a cutoff prints a note on stderr).
| Column | Type | Notes |
|---|---|---|
Timestamp |
DateTime64(9) | When the log was emitted; supports time arithmetic (Timestamp > now() - INTERVAL 1 HOUR). |
ServiceName |
String | Service that emitted the log. |
SeverityText |
String | |
SeverityNumber |
UInt8 | |
Body |
String | The log message. |
LogAttributes |
Map(String, String) | Per-event attributes. |
ResourceAttributes |
Map(String, String) | Resource labels (for example, service.name, deployment.environment.name). |
TraceId |
String | Distributed-tracing trace ID, if present. |
SpanId |
String | Distributed-tracing span ID, if present. |
Also exposed but rarely queried directly: EnvironmentId, ProjectId (used for auto-scoping), TraceFlags.
observability.metrics
Unified metrics from services. Retention: up to 90 days.
| Column | Type | Notes |
|---|---|---|
TimeUnix |
DateTime64(9) | Measurement timestamp; supports time arithmetic. |
MetricName |
String | For example, container.cpu.time (see below). |
MetricType |
String | 'gauge' or 'sum'. |
Value |
Float64 | Measured value. |
ServiceName |
String | Service that emitted the metric. |
Attributes |
Map(String, String) | Metric dimensions. |
ResourceAttributes |
Map(String, String) | Resource labels. |
Service metrics
Emitted for every running service container:
| Metric | Type | Notes |
|---|---|---|
container.cpu.time |
sum | Cumulative CPU seconds; rate it for utilization. |
container.cpu.usage |
gauge | Current CPU usage in cores. |
container.cpu.limit_utilization |
gauge | Fraction of CPU limit used (0–1). |
container.memory.working_set |
gauge | Memory in active use (bytes), the right number for “is this service close to OOM”. |
container.memory.usage |
gauge | Total memory usage including page cache (bytes). |
container.memory.available |
gauge | Memory available before the container limit (bytes). |
k8s.volume.available |
gauge | Free space on attached volumes (bytes). |
k8s.volume.capacity |
gauge | Total capacity of attached volumes (bytes). |
observability.traces
OpenTelemetry spans. Today every row is the edge span the platform’s ingress records for a request to one of your public services (Source = 'ingress'): one span per request, so counting rows counts requests. Requests between your services never pass the ingress and are not represented; neither are responses served from the CDN edge cache. Retention: up to 30 days.
| Column | Type | Notes |
|---|---|---|
Timestamp |
DateTime64(9) | Span start; supports time arithmetic. |
TraceId |
String | Trace ID. The ingress forwards it upstream as a W3C traceparent header, so an app that logs it can be correlated with observability.logs.TraceId. |
SpanId, ParentSpanId |
String | ParentSpanId is empty for the edge span. |
SpanName, SpanKind |
String | SpanKind is 'Server' for ingress spans. |
ServiceName |
String | Your service that handled the request. |
SpanAttributes |
Map(String, String) | HTTP details, see below. |
ResourceAttributes |
Map(String, String) | service.name, specific.project.id, specific.environment.id. |
Duration |
UInt64 | Nanoseconds, as OpenTelemetry defines it; divide by 1e6 for milliseconds. |
StatusCode, StatusMessage |
String | The OpenTelemetry span status ('Unset', 'Ok', 'Error'), not the HTTP status. |
Source |
String | 'ingress'. |
SpanAttributes keys (values are strings):
| Key | Notes |
|---|---|
http.method |
GET, POST, … |
http.target |
Request path including the query string. |
http.route |
Path with ID-looking segments replaced by {id} (/api/users/{id}) and static-asset file names by {file}.ext (/assets/{file}.svg), a heuristic since the platform has no route manifest. Group by this for per-endpoint views. |
http.status_code |
HTTP status as a string: toUInt16OrZero(SpanAttributes['http.status_code']) >= 500. |
http.server_name |
Hostname the request arrived on. |
http.user_agent, http.referer, http.x_forwarded_for, net.peer.ip |
Client details. |
nginx.upstream_header_time |
Seconds until your service sent its response headers (time to first byte), i.e. how fast the service answered. Unaffected by streamed responses and slow client downloads, so use this for “is my service slow”. A comma-separated list when nginx retried another pod; absent when the request never reached your service (nginx’s own 408/499). |
nginx.upstream_response_time |
Seconds until your service finished sending its response. Duration additionally runs until the client received the last byte. Same list/absent conventions as above. |
observability.events
Kubernetes events about your service pods (probe failures, restarts, out-of-memory kills), in the OpenTelemetry log-record shape. Retention: up to 30 days.
| Column | Type | Notes |
|---|---|---|
Timestamp |
DateTime64(9) | When the event was recorded. |
ServiceName |
String | Your service the pod belongs to. |
SeverityText, SeverityNumber |
Kubernetes Warning events, plus the Normal “Killing” event that marks a pod shutting down. |
|
Body |
String | The event message, e.g. Readiness probe failed: .... |
LogAttributes |
Map(String, String) | k8s.event.reason (Unhealthy, BackOff, OOMKilled, Killing, …), k8s.event.count, k8s.event.action. |
ResourceAttributes |
Map(String, String) | k8s.pod.name, k8s.object.kind, k8s.object.name, service.name, specific.*. |
observability.alert_evaluations
Every evaluation of your default alerts. This is what the alert graph in the dashboard shows; query it to see the exact numbers that did or didn’t fire. Retention: 30 days.
| Column | Type | Notes |
|---|---|---|
Timestamp |
DateTime64(3) | When the alert was evaluated. |
AlertName |
String | http-5xx, http-p99-latency, probe-failures, restart-loops. |
ServiceName |
String | The service the alert is about. |
Outcome |
String | value, no_data, error (the query failed), or paused (skipped during a volume-backed rollout). |
Value |
Nullable(Float64) | The alert query’s result; NULL unless Outcome is value. |
State |
String | The alert’s state after the evaluation: ok, firing, error. |
Default alerts
Every deployed service gets a set of default alerts, evaluated every minute from these same views and listed under Alerts in the dashboard; the project health indicator reflects whatever is firing. Each alert is a plain SQL query returning one number, compared against a threshold: open it in the dashboard to see the exact query, and paste it into specific query to get the same value the alert saw. What the alerts measure, how slow responses are judged, and how notifications work is covered in the Alerts guide.
Debugging recipes
Recent errors for a service:
SELECT Timestamp, Body
FROM observability.logs
WHERE ServiceName = 'api'
AND SeverityNumber >= 17
AND Timestamp >= now() - INTERVAL 1 HOUR
ORDER BY Timestamp DESC
LIMIT 50
Search log bodies for a substring:
SELECT Timestamp, ServiceName, Body
FROM observability.logs
WHERE positionCaseInsensitive(Body, 'connection refused') > 0
AND Timestamp >= now() - INTERVAL 6 HOUR
ORDER BY Timestamp DESC
LIMIT 100
Service CPU timeseries in 1-minute buckets:
SELECT
toStartOfMinute(TimeUnix) AS bucket,
avg(Value) AS cpu_cores
FROM observability.metrics
WHERE MetricName = 'container.cpu.usage'
AND ServiceName = 'api'
AND TimeUnix >= now() - INTERVAL 1 HOUR
GROUP BY bucket
ORDER BY bucket
Correlate all logs belonging to one trace:
SELECT Timestamp, ServiceName, Body
FROM observability.logs
WHERE TraceId = '<trace-id>'
ORDER BY Timestamp
Requests, error count and p95 latency per endpoint over the last hour:
SELECT
ServiceName,
SpanAttributes['http.method'] AS method,
SpanAttributes['http.route'] AS route,
count() AS requests,
countIf(toUInt16OrZero(SpanAttributes['http.status_code']) >= 500) AS errors,
round(quantile(0.95)(Duration) / 1e6) AS p95_ms
FROM observability.traces
WHERE Source = 'ingress'
AND Timestamp > now() - INTERVAL 1 HOUR
GROUP BY ServiceName, method, route
ORDER BY requests DESC
LIMIT 50
Performance
The views are partitioned and ordered to make environment-scoped, time-bounded queries fast. To stay within the 30-second limit:
- Always filter by time -
Timestamp >=for logs, traces and events,TimeUnix >=for metrics. Unbounded scans across 30/90 days will time out. - Add
ServiceNameand/orMetricNameto narrow the partition whenever you can. - Select only the columns you need -
Body,LogAttributes, andResourceAttributesare the heavyweight columns. LIMITexploratory queries while shaping them.
Discovering data
When you don’t know what’s there, ask the database:
SHOW TABLES FROM observability;
DESCRIBE observability.logs;
SELECT DISTINCT ServiceName FROM observability.logs LIMIT 50;
SELECT DISTINCT MetricName FROM observability.metrics LIMIT 100;
ClickHouse dialect
The observability store is ClickHouse, so specific query without --db uses the ClickHouse SQL dialect (unlike --db, which is standard Postgres). Commonly needed functions:
now(),now64()- current timetoStartOfMinute(t),toStartOfFiveMinutes(t),toStartOfHour(t)- bucket timestampspositionCaseInsensitive(haystack, needle)- case-insensitive substring searchJSONExtractString(s, 'field'),JSONExtractInt,JSONExtractFloat- pull fields out of JSON log bodiesLogAttributes['key'],Attributes['key'],ResourceAttributes['key']- read map columns