docsOperate

Minimum Requirements (running beyond local)

This page is the short answer to "what do I actually need to stand memQL up somewhere real?" — i.e. anything past the local dev cluster. It states the minimums and links the deep runbooks for each piece. The biggest non-obvious requirement is the database, so that section is first and most detailed.

There are two run modes:

ModeWhat it isHow
Local devk3d + ArgoCD cluster on one machine — Postgres, the mesh, LiveKit. Throwaway.make up (primary local path, memql#2061). See Reproduce staging locally.
Real deployment (staging / prod / self-host)A Kubernetes mesh against a managed TimescaleDB.The rest of this page.

Everything below is environment-agnostic by design: the architecture is identical local → staging → prod; only the config values (DSNs, replica counts, sizes) differ. If a requirement is "do X differently in prod," that's a bug, not a feature.


1. Database — Tiger Cloud is the only supported provider today

memQL stores everything (the time-series memory graph + the observability hypertables) in PostgreSQL 16 + the TimescaleDB extension. For a real deployment that means a managed Tiger Cloud service (TIMESCALEDB type). Tiger Cloud is the only DB provider we support right now — the engine assumes TimescaleDB (hypertables + continuous aggregates back code_invocation observability), and the deploy tooling, connection model, and runbooks are all written against Tiger. A vanilla self-managed Postgres is not a supported target yet.

Minimum DB checklist

  • A Tiger Cloud service, PostgreSQL 16 + TimescaleDB.
  • Sized so the connection budget fits the mesh (see below) — start at 1 CPU / 4 GB and max_connections = 500 for a small staging mesh.
  • The transaction pooler enabled (PgBouncer, transaction mode).
  • Two connection strings wired into memql-secrets: MEMQL_DATABASE_DSN → the pooler, and MEMORY_NODES_DATABASE_DIRECT_DSN → the direct endpoint.

Why two endpoints? — the connection model (epic #1925)

The whole mesh (≈10 node-types × replicas) shares one database. Tiger caps direct max_connections per tier (25–500 on the 0.5–4 CPU range) and reserves ~17 for superuser/ops, so the direct budget is small and fixed — and a deploy surge (blue-green + rolling restart briefly doubles pods) used to blow past it → SQLSTATE 53300 ("remaining connection slots…") storms that wedged a roll. Tiger's intended answer to "many connections" is the connection pooler, not a bigger max_connections.

So memQL runs a hybrid endpoint split:

  • Bulk traffic (all queries + mutations, every pod) rides MEMQL_DATABASE_DSN → the transaction pooler. Client connections decouple from Postgres backends, so a deploy surge no longer maps 1:1 to slots (transaction-pool ceiling ≈ (max_connections − 17) × 20).
  • Session-stateful work — session-scoped advisory locks (cognition dispatch/greet/feedback gates, cron leader, topology reconciler, planner admission) and migrations — rides MEMORY_NODES_DATABASE_DIRECT_DSN → the direct endpoint. A transaction-mode pooler recycles a server backend between statements, which would silently drop a held session lock; these few, bounded connections take a real slot instead.

In code: Database.DirectBunDB() returns the direct pool when DIRECT_DSN is set, else falls back to the main pool — so local/dev without a pooler is unaffected (single pool, identical behaviour). bun's pgdriver speaks the simple query protocol (no server-side prepared statements), so transaction pooling is safe.

Sizing the budget

text
peak_direct ≈ Σ(session-stateful holders) + migrate(1) + live FOREIGN backends
REQUIRE: peak_direct ≤ max_connections − reserved(~17)
  • Foreign backends are real and must be budgeted. Tiger's own control-plane process (application_name=deployer, the postgres superuser) holds a pool of connections you cannot terminate (#1822). Size max_connections with headroom above it (this is why staging runs 500, not 105).
  • Bulk pods do not count against the direct budget — they multiplex through the pooler.

Full detail, the budget formula, the pre-deploy gate, and the monitor: DB connection budget & graceful deploy. Tiger CLI / service management: Database Setup. max_connections is set in the Tiger console → Common parameters (not the CLI/SQL).


2. Compute — Kubernetes

  • A Kubernetes cluster. We run and test on Azure Kubernetes Service (aks-memql-staging); any conformant cluster should work, but AKS is the exercised path. See Infrastructure.
  • No database pod — the DB is the managed Tiger service above.
  • The engine mesh node-types (each a Deployment, 2 replicas for HA): identity, cognition, voice, agent, planner, workbench, mcp, plus livekit and the voice-agent. Manifests: deploy/k8s/base. A product stack adds a bff -- a plain, product-agnostic engine node that fronts the product's DSL bundle -- plus the product client (SPA), both layered in from the product's own overlay -- see Downstream product stacks.
  • Rough small-staging footprint: ~4 × 2-vCPU nodes. Right-size per workload.
  • Argo CD for GitOps (+ Argo Rollouts when a product stack runs a blue-green bff; the controller install lives in deploy/rollouts/install).

3. Secrets & auth

Every DB-connecting pod mounts the memql-secrets Secret via envFrom. The four keys (see deploy/k8s/base/secret.example.yaml):

KeyWhat
MEMQL_MASTER_KEY32-byte key that decrypts the genesis envelope.
MEMQL_GENESIS_B64base64 of the sealed env envelope (~150 config vars, decrypted in-process at boot; MEMQL_GENESIS_AUTOLOAD=true).
MEMQL_DATABASE_DSNDB — the transaction pooler endpoint.
MEMORY_NODES_DATABASE_DIRECT_DSNDB — the direct endpoint.
  • Identity is the in-house auth service; for multi-replica HA it needs a shared MEMQL_IDENTITY_SIGNING_KEY_B64 (Ed25519 seed) in the envelope — every replica derives the same key/JWKS. See Identity Service and the Access Model.
  • AI providers: an MEMQL_OPENAI_API_KEY is required (cascade voice = OpenAI ASR/TTS); Anthropic optional. See Environment Variables for the full env surface and the bootstrap-envelope vs concept-stored split.

4. Images — built on the build server, never a laptop

Deployable images are built on GitHub Actions → OIDC → ACR acrmemql, never hand-built locally (local Docker is dev-only):

  • memql (this repo) → every node-type image — identity, bff, cognition, agent, planner, voice, workbench, mcp — as a product-agnostic engine image (.github/workflows/build-engine-images.yml). There are no per-product node images: a plain engine node runs a product's DSL at runtime, so it does not need product code compiled in and does not fail on a product prompt template — the DSL bundle supplies the prompts (next bullet).
  • The product → a tiny data-only DSL bundle image (just the .memql tree), mounted at MEMQL_DSL_PATH by the deploy/k8s/components/dsl-bundle init-container so every engine node loads the product's domains at boot.
  • The product frontend repo → the client (SPA) image.

So a product ships just two artifacts — a DSL bundle and a client — and rides the same product-agnostic engine images as everyone else. A release is {engine version, bundle digest, client digest} pinned in one per-env overlay (see Deploy, below) — there is no separate carrier build workflow or release lockfile. See Downstream product stacks for the contract.


5. Deploy — GitOps, digest-pinned, connection-gated

  • The per-env overlay (deploy/k8s/overlays/<env>) is the single image authority — every image pinned by @sha256: digest (a CI gate enforces it). Argo CD reconciles the overlay; rollback = git revert.
  • Migrations run once (a gated pre-deploy migrate Job + identity on boot), on the direct endpoint.
  • Connection safety is a deploy requirement, not an afterthought (#1958):
    • Pre-deploy gate scripts/deploy/conn-headroom-check.sh --live — blocks a deploy whose projected peak + live foreign backends would exceed the budget (catches "deploy into an already-full instance").
    • */5 monitor (conn-monitor CronJob) — logs total-vs-budget + a per-application_name leak detector, so pressure is seen before it storms.
  • Cutover ordering gotcha: an Argo CD image-sync re-applies a Deployment's replicas past ignoreDifferences, so it un-drains a scaled-to-0 fleet. When changing a connection-config secret, cut the secret over before the sync scales pods up; bring identity up first. See DB connection budget and Deployment Strategy (see the product pack repo's docs/operate/deployment-strategy.md).

6. Versions

ComponentMinimum
Go1.26.1+ (to build)
PostgreSQL16 + TimescaleDB (Tiger Cloud)
Kubernetesa recent conformant cluster (AKS is the exercised target)
Argo CD / Rolloutsrequired for the GitOps + blue-green path

See also