Replaces the pseudo-version pin (v1.49.2-0.20260719024505-24ff20a68f52, the
Bug-A iterator-leak fix only) with the released v1.49.2, which also carries the
Bug-B guard: org.Resolve skips the doomed GetById for a legacy all-digit cached
id (IAM Valkey's stale 1772587477 for 'hanzo') and resolves by name, so the
(*Query).ById legacy-numeric path that hot-looped in v1.801.95 is never taken.
Keeps ai v1.826.4 (in-proc TierReader). go.mod+go.sum only.
The auth fix let the metering cap check actually reach commerce AuthorizeSpendCap; a
legacy-org GetById hot-loop there then HUNG every completion (no timeout on the
in-proc authorize) — a cap that can block/hang the completion path is worse than one
that does not enforce. scopeAuthorize now runs the authorize under a strict 1.5s
deadline AND a select-based hard timeout that returns even if the in-proc handler
goroutine is STUCK (an unresponsive hot-loop cannot be interrupted, so ctx alone would
not unblock). On timeout OR any error -> AuthorizeVerdict fails OPEN (allow) — a slow,
broken, or hot-looping commerce ALWAYS allows, never waits. OnCapError logs each
fail-open so a degraded cap is observable. Regression test: a 10s-hanging authorize
returns an ALLOW in ~1.5s (completion never hangs).
The commerce hot-loop itself (the root cause) is fixed separately; this timeout is the
non-negotiable safety net that makes the cap path unable to hang regardless.
v1.826.4 ships the LLM-as-judge dense-reward loop fully activated: the Mean-Field
Judge Panel (diverse calibrated judges, reputation-weighted consensus), geo-aware
consent (EU/UK/EEA explicit opt-in via CF-IPCountry, non-EU opt-out default), judge
config dynamic at admin.hanzo.ai (OrgSettings "*" row, no env), MFJP enabled by
default on a diverse cheap panel, and internal dev orgs seeded on. Judge scoring uses
the existing probe service bearer (no new secret). Also carries the MFJP + scientific
proof from v1.826.3.
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
commerce 24ff20a6 closes single-row query iterators (Query.First). This
stops the Postgres pool leak that starved org.Resolve on the co-resident
balance + per-tier gate path — the 'context deadline exceeded' that made
the Enso per-tier SKU gate fail open and spiked chat latency to 10-40s.
Wire() gained the /v1/dns zone plane (after projects) and /v1/cloudflare edge
plane (after integrations) but the frozen golden in wire_test.go was not
updated, so TestWireOrderMatchesFrozen failed (87 specs vs 85 frozen) — which
red-lit cloud's CI/CD and blocked the auto-release image build. Refreeze the
golden to the exact runtime sequence (verified position-by-position, 87==87).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The ONE Kubernetes noun, proxied to Visor (clients/visor/k8s.go): list the
org's DOKS clusters, one cluster's detail (node pools + worker nodes), DEPLOY
(create) / delete clusters, and the fleet-wide worker nodes. Reads are org-scoped
by the validated IAM owner; mutations (create/delete) are admin-gated
(principal.IsSuperAdmin || IsOrgAdmin) — real house-account infra spend.
Consolidates the worker-node consumption: managedMachines now reads
/v1/k8s/nodes (was /v1/kubernetes-nodes), matching Visor's consolidated path —
no parallel kubernetes-* surface remains.
- k8s.go: listK8sClusters / getK8sCluster / createK8sCluster (admin) /
deleteK8sCluster (admin) / listK8sNodes; wire structs + view mappers.
- visor.go: mount the /v1/k8s/* group; managedMachines -> /v1/k8s/nodes.
- tests: proxy + tenant-scoping, detail shape, nodes, and the admin gate
(a non-admin create/delete is refused BEFORE reaching Visor); the fleet
DOKS-node fake tracks the new /v1/k8s/nodes path.
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
- buildJobSpec doc claimed the REVERTED over-hardening (allowPrivilegeEscalation=
false, all caps dropped); correct it to the actual documented rootless posture
(defaults left for rootlesskit newuidmap) + point at the securityContext.
- tenantPullSecretName comment overstated "cloud-api holds no secrets grant";
clarify cloud's only Secrets write is the per-tenant KMS-auth creds in a TENANT
ns, and that the isolated build ns must stay OFF the tenant-RBAC selector so no
secrets grant is projected there (R6).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
GET /v1/platform/projects 500'd on console dashboard init. The iamStore guard
converts a nil co-resident IAM object store into a typed 503, but listProjects
re-stamped ANY store error as a 500 (zip.Errorf(500, "list: %v", err)),
discarding the status — so a signed-in session's first read broke dashboard
init with {"status":500,"error":"list: platform requires the co-resident IAM
store, which is not initialized"}.
The dashboard's first authenticated read now degrades any store failure to an
empty project set (200 []) — a new org genuinely has zero projects — logging the
real cause for operators (never swallowed), written in-band so no outer error
filter can reflatten it. Also guards a stray nil row from nil-derefing into a
500. The store-level 503 guard + its three unit tests are unchanged.
Repro + regression gate: TestListProjects_NilIAMStore_ServesEmpty200 (real
iamProjects over a nil in-process IAM engine — the deployed condition) and
TestListProjects_StoreError_ServesEmpty200.
RED found the /v1/paas auth broadening handed every brand-org ("hanzo") OrgAdmin
fleet-wide rolling-restart of the platform's OWN tier (the only namespaces the board
scans are hanzo{,-testnet,-devnet}, where iam/kms/gateway/cloud/… run) — a live DoS
lever, partially re-opening the 2026-07-08 admin-org P0.
H1: the MUTATING POST /v1/paas/apps/:app/deploy now uses operatorGuard (principal.
IsSuperAdmin ONLY), not the broad read guard. Restarting a shared platform service is
a platform-operator action; a customer-org admin — even of the brand org — is refused
403. The READ board (list/get) stays SuperAdmin||OrgAdmin (observe, audit-logged,
bounded). Confinement (scopedNamespaces) unchanged.
L1: deploy REQUIRES ?env=main|test|dev (nsForEnv-validated) — a bare deploy no longer
silently targets production; the CLI requires --env before the call.
Tests: TestDeploy_OrgAdmin_403_Platform (the H1 regression), _NonAdmin_403,
_SuperAdmin_RollingRestart, _RequiresExplicitEnv, _SuperAdmin_EnvSelectsNamespace;
CLI TestDeployRequiresEnv. All green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The embedded ai per-tier SKU gate (family_tier.go) was fail-open in-cluster: it
resolved the caller's tier with an authed HTTP self-call to the cloud edge, which
401/403s a service token on /v1/billing/*, so the gate saw "" and admitted every
tier — enso/enso-ultra were open to free callers.
Mirror wireFinance's SetBalanceReader: install aiobject.SetTierReader so ai reads
the subscription tier DIRECTLY over the co-resident commerce client the metering
gate already bills over (commerceinproc in-process, with the service token commerce
itself accepts) — never the cloud edge. Add metering.Client.Tier to decode tier.name
from GET /v1/billing/tier. Fail-safe preserved: a commerce error or unknown tier
folds to "" (allow), so a commerce blip never locks out a paying caller.
Bumps ai v1.824.2 -> v1.825.2 (the object.TierReader seam).
managedMachines unioned Visor's registry (/v1/get-machines) + live droplet list
(/v1/machines); a DOKS cluster's worker NODES appeared in neither (their droplet
carries a k8s tag, not a hanzo-org droplet tag), so world.hanzo.ai showed
standalone droplets but never cluster nodes.
Add GET /v1/kubernetes-nodes as the THIRD source (Visor unions the house-account
hanzo-org-tagged clusters + BYOC Provider.ClusterID clusters and returns each
worker node as a Machine keyed by droplet id). It is processed after registry and
live, so a DOKS node whose droplet is ALSO in the live list dedupes by droplet id
and never lists twice; a cluster-only node surfaces. Independently resilient like
the other two — a kubernetes-nodes outage is logged and skipped, never hiding the
registry/live/BYO sources.
Test: TestMachinesMergeDOKSNodes — a DOKS-only node appears, and a node whose
droplet is already live collapses BY ID (the node row carries a different name, so
only id-dedup can merge it). Full clients/visor suite green (31 subtests).
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
POST /v1/event is the ONE ingestion door: body is a single Event or a JSON-array
batch (no /v1/event/batch), org resolved IAM-only and fail-closed (eventTenant),
funneled through the ONE write core (ingestEvents) into hanzo.events. The
Segment/beacon (/v1/analytics,/v1/tracker) and PostHog (/v1/insights/e) wires
become thin DEPRECATED adapters over the same core. Org is never read from body.
capture resolves a presented project/API key to its owner org through the single
IAM key resolver (sharedKeys, 60s cache incl. miss-cache). A presented-but-
unresolvable key is refused (403) and NEVER falls through to the brand-host
fallback, so a keyed request can never cross-tenant write. Anonymous marketing
traffic still resolves to the public brand org server-side from Host.
Brings the merged Enso router fixes into the deployed cloud binary:
- #107 (v1.825.0): auto never routes to a family SKU it can't serve + forward the
resolved model (withModel body rewrite) → model=auto serves 200 (was 404); grant-
aware known predicate; flywheel boots from the single shared Bootstrap.
- #108 (v1.825.1): trainer fits EARLY (~90s after boot) then cadence → completed
retrain cycles survive frequent redeploys (churn-resilient).
./apps (ai.Mount site) compiles clean against v1.825.1 (API-compatible).
Co-authored-by: zeekay <zeekay@hanzo.ai>
The metering cap-gate authorize (and the SuperAdmin cap-oversight Forward) call the
in-proc /v1/billing/spend-alerts/authorize with the COMMERCE_SERVICE_TOKEN, but the
customer /v1/billing/* bridge required a validated IAM principal -> "sign in to view
billing" (403) -> the cap fails-open and never enforces. billingData now admits a
trusted S2S caller carrying the verified service token: org from the EdgeAuth-controlled
X-Org-Id, query forwarded VERBATIM (a trusted caller names its own subject), no
subject-pin. A public caller can never present it (the gateway 401s a Bearer that is
not an IAM JWT / hk-|pk-|sk- key), and an unauthenticated caller still gets 403 (tested).
Adds spend-alerts/authorize to the GET allowlist; constant-time token compare.
ATOMIC with the flag: bumps commerce to v1.49.2 (SPEND_CAP_ENFORCE, default OFF), so
the instant the cap can reach the handler the enforcement gate is fail-open in the
binary -> auto-deploy stays safe until an operator flips the flag after the canary proof.
Security invariants tested: public/unauth -> 403; wrong bearer -> 403; verified token +
X-Org-Id -> forwards verbatim; token без X-Org-Id -> 403.
The per-org Cloudflare asset plane (Pages/Workers/R2/KV/D1) is repointed from a
first-level /v1/cloudflare/* surface to /v1/integrations/cloudflare/*, so a
third-party provider is connected AND used under one unified namespace — matching
where the connector's connect/callback/verify/disconnect legs and the KMS token
coordinate already live. Pure path move: no auth, isolation, or handler logic
changes. Route registrations, doc comments, and tests repointed together.
No collision with the connector's parametric routes: the asset paths are all
3+ segments (/cloudflare/{pages,workers,r2,kv,d1}/...) while the connector's
/v1/integrations/:provider and /:provider/{connect,callback,disconnect,verify}
are 1- and 2-segment patterns whose literal second segment never equals an asset
group. Static-under-param co-registration is already proven in the integrations
plane (slack/link static beside :provider). Mount order unchanged: integrations
before cloudflare.
Assisted-by: Claude:claude-opus-4-8
The first rootless spec over-hardened (allowPrivilegeEscalation:false +
capabilities drop ALL), which breaks rootlesskit's setuid newuidmap/newgidmap
sub-uid mapping — proven by an on-cluster canary:
newuidmap ... failed: operation not permitted
Relax to the documented moby/buildkit k8s rootless posture: privileged:false,
runAsUser/Group 1000, runAsNonRoot, seccomp+AppArmor Unconfined, and leave
allowPrivilegeEscalation / default caps at k8s defaults (newuidmap needs them).
Still user-namespaced, no host root — the decisive win over privileged=true.
Re-canaried: rootless build + scoped push-hanzoai cred pushed to ghcr OK.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
`hanzo apps list`, `hanzo deploy`, `hanzo clusters` targeted the OLD TS-Dokploy
contract (/v1/apps, /v1/org/{org}/cluster, /v1/org/.../redeploy) — all 404 on the
live Go cloud (ghcr.io/hanzoai/cloud). Repoint the CLI at the endpoints the Go
cloud actually serves, authorized off the SAME IAM login `hanzo build` now uses
(no --platform-token). Drift confirmed live as z@hanzo.ai:
/v1/apps → 404 /v1/paas/apps → 403 (was SuperAdmin-only)
/v1/org/*/cluster → 404 /v1/clusters → 200 (already org-scoped)
/v1/platform/projects → 500 (co-resident IAM off)
CLI (cli/platform.go, cli/commands.go):
apps list/get → GET /v1/paas/apps[/{app}] (no client org filter — the board is
confined to the caller's org SERVER-side by the validated identity)
deploy <app> → POST /v1/paas/apps/{app}/deploy (rolling restart; --env selects
the lifecycle namespace; org from identity, not the path)
clusters list/get → GET /v1/clusters (Visor-managed + BYO; org from identity)
Removed the TS-contract vestiges with NO Go backend: `apps sync` (the board is
live-computed), `clusters create/select/install-baseline/target` and `k8s target`
(DOKS provisioning + deploy-target selection are not implemented on the Go cloud).
Reshaped the Cluster DTO to the live visor clusterView (dropped dead Phase/Active/
operator/baseline fields). Platform client doc corrected: it CAN validate IAM tokens.
Backend (clients/paas): authorize the fleet board off ONE IAM identity, exactly like
/v1/runner (clients/platform/runner.go). guard now admits a validated principal who
is a SuperAdmin OR an OrgAdmin (principal.Validated + IsSuperAdmin || IsOrgAdmin —
the ONE verifier, unforgeable off-gateway), and each handler CONFINES a non-super
caller to the platform namespaces its own validated org owns (scopedNamespaces, keyed
on principal.Org — never a client header): a SuperAdmin sees the whole fleet, an
OrgAdmin only its own org (empty board / clean 404 otherwise), so a tenant admin can
never observe — or restart — another org's, or a platform, app. `?org=` cannot widen
the view (confinement is at the namespace scan, before the filter).
deploy is now a real zero-downtime ROLLING RESTART (the kubectl-rollout-restart
mechanism: stamp the pod-template hanzo.ai/restartedAt annotation) instead of the
409-refuse. It never changes the declared TAG (that stays a git commit, the one thing
Hanzo CD's selfHeal reconciles), so there is no drift to revert — the honest,
GitOps-compatible "redeploy this app". resolveTarget split so the machine release
path (release.go) keeps the full scan; the identity-scoped deploy/getApp use the
caller's authorized namespaces.
SECURITY: this broadens auth on a control-plane surface (new mutating restart path).
Mirrors blue's IAM-admin pattern; flag for red. TDD: 12 new paas cases (role gate,
tenant confinement on list/get/deploy, forged-org cannot widen, rolling-restart lands
the annotation, foreign-org 404 + no mutation) + the CLI path/DTO tests, all green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
RED re-review of the unify-infra PaaS-auth flip found 2 HIGH + 1 MED on the
privileged build endpoint. Fixes:
H1 (borders CRITICAL) — cross-org supply-chain push. imageAllowed() permitted
ghcr.io/{hanzoai,luxfi,zooai}/ regardless of the caller's validated org, so any
org-admin could overwrite another brand's prod image via the shared push cred.
Bind the image's registry-org to the caller's org (orgRegistryNamespaces map);
only a real SuperAdmin may cross. The machine (fabric) token keeps full
owned-registry latitude. Cross-org now 403; same-org 202.
H2 — privileged rootful buildkit in the main platform namespace with the shared
3-org push cred. buildJobSpec is now ROOTLESS (moby/buildkit:*-rootless, uid 1000,
no privileged, no privilege-escalation, caps dropped, --oci-worker-no-process-
sandbox), runs in a DEDICATED isolated namespace (CLOUD_PLATFORM_BUILD_NS default
→ hanzo-build, off the platform ns), and mounts ONLY the target org's push
credential (push-<namespace>), never the shared kaniko-ghcr. Node-pool taint +
automountServiceAccountToken:false retained.
M1 — `image` bypassed validateBuildInputs → buildkit --output attribute
injection (ghcr.io/hanzoai/x,registry.insecure=true). Added validateImageRef
(strict single OCI ref, rejects comma/space/quote/newline/'='), folded into
validateBuildInputs and enforced early at the handler.
I2 — the identity validator now fail-closes on an empty resolved issuer OR
audience set instead of silently disabling that axis; issuerAllowed denies on
an empty trusted set.
TDD: cross-org 403 + same-org 202 + SuperAdmin cross-org + orgless-403 +
image-injection-reject + rootless/scoped-cred spec + empty-trust-set-deny all
green (pure-Go, as prod ships).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The prior prefix guard checked fasthttp's URI().Path(), which decodes only ONE
layer of percent-encoding. A DOUBLE-encoded traversal survives that one decode as
a literal %2e/%2f that still KEEPS the /v1/dns/ prefix -- so the prefix check
passes, cloud forwards base + /v1/dns/%2e%2e/admin, and the upstream decodes the
second layer to /v1/admin. Proven bypasses: /v1/dns/%252e%252e/admin,
/v1/dns/%252e%252e%252fadmin, /v1/dns/..%252fadmin.
After the prefix check, also refuse any once-decoded path that still carries a `%`
(a still-encoded byte => the client double-encoded) or `..` (residual traversal).
Neither appears in a legitimate DNS-API path -- zone labels are DNS names /
punycode xn--, and the query string (checked separately) is unaffected. Fail
closed 400 before a byte leaves cloud.
Also relay the upstream Location header so a 3xx -- never followed, per
CheckRedirect -- passes back verbatim (status + Location) as the comment claims,
rather than being silently dropped.
Regression: the escaped-path test gains the 3 double-encoded vectors (each refused
400 with 0 upstream bytes), plus a redirect test proving an upstream 302 is not
followed and its Location relays verbatim. 9 tests green.
Assisted-by: Claude:claude-opus-4-8
The forward head built the upstream target from uri.Path(), which is NORMALIZED
and percent-decoded. Fiber matches the /v1/dns/* wildcard on the RAW path, so a
dot-segment or encoded-dot traversal (/v1/dns/../../admin/secrets,
/v1/dns/..%2f..%2fadmin, /v1/dns/../../../metrics) still routed to the handler
while the normalized path escaped the prefix -- letting the caller drive the
WHOLE path on the DNS host. Contained today only because the upstream 404s
unknown paths; a latent path-scope escape the moment :8443 serves anything else.
Guard the normalized path fail-closed BEFORE building the target: require it to
be exactly /v1/dns or under /v1/dns/, else 400 and forward nothing. Because the
path is already normalized, every traversal/encoded-dot escape fails this check.
Correct the comment that wrongly claimed the path was locked by the route match
(the host-pinning claim was, and stays, true).
Also stop the shared http.Client from following upstream 3xx (CheckRedirect =>
http.ErrUseLastResponse) so redirect responses pass through verbatim and a 3xx
can never silently re-target the request onto another host or path.
Regression test proves fail-closed: each escaped path is refused (400) and 0
bytes reach the upstream. Existing 7 tests stay green (8 total).
Assisted-by: Claude:claude-opus-4-8
Addresses Red's FIX-THEN-SHIP findings:
- HIGH: stamp X-Hanzo-Org (the served org) on every /v1/cloudflare response so a
per-org caller can detect a pinned/comingled org. The platform Pages client asserts
it equals the requested org and fails LOUD if a non-org-switch-capable service token
made the identity boundary pin X-Org-Id to the token's own owner — no silent
cross-tenant read/write.
- MEDIUM: gate mutations (POST/PUT/DELETE) on principal.IsOrgAdmin via a new authWrite
front door (reads stay validated-org-only), parity with the AdminOnly connector. A
non-admin is refused before any KMS token read and never reaches Cloudflare.
- LOW: resolveAccount now prefers the account captured at connect time
(integrations.ConnectionFor ExternalID), falling back to live /accounts discovery
only when none is stored — no per-call round-trip, deterministic for multi-account
tokens.
Tests: +TestResponseStampsActingOrg, +TestMutationRequiresOrgAdmin,
+TestStoredAccountSkipsDiscovery; existing mutation tests drive as org admin. 12/12
pass, -race clean.
Assisted-by: Claude:claude-opus-4-8
New cloud subsystem clients/cloudflare exposing /v1/cloudflare/{pages,workers,r2,kv,d1}/*,
sibling to hanzodns's /v1/dns. It reads each org's KMS-sealed Cloudflare token in-process
through the integrations custody seam (integrations.TokenFor) and proxies to the Cloudflare
API v4 with the cfDo shape reused verbatim from hanzodns — no global env token, no
bearer-relay hop (that is only hanzodns's separate-process need).
Tenant isolation: org is derived ONLY from the validated principal (principal.Org); the KMS
token path is keyed on that org, so cross-org token reach is structurally impossible and an
unvalidated request fails closed (403). Pages (project CRUD, deploy, custom-domain add/delete)
and Workers (script put/list/delete via multipart module upload, workers.dev subdomain, zone
route bind/list/delete) are wired; R2/KV/D1 ship typed provider methods + routes that answer
an honest 501 (never a fake success).
Appends the Workers connector scopes (Account:Workers Scripts:Edit, Zone:Workers Routes:Edit)
for the now-callable capabilities and wires the subsystem into apps.Wire after integrations.
Assisted-by: Claude:claude-opus-4-8
House rule: no extraneous /api/, just /v1/. The projection API moves from
/v1/deploy/api/v1/* → /v1/deploy/<resource> (settings, session/userinfo, version,
account/can-i, applications, applications/{name}/resource-tree, .../{sync,rollback}).
This IS the deploy API now — the superseded native /v1/deploy/{applications,:name/*}
routes are removed (their readers stay, reused by the projection). health +
reconcile unchanged. Guard test updated.
console.hanzo.ai serves the DnsModule but cloud held no /v1/dns head, so
console.hanzo.ai/v1/dns/* 404'd and the dashboard showed empty zones. Add a thin
forward head (clients/dns) that relays each /v1/dns/* request to the DNS control
plane (HANZO_DNS_URL, default the in-cluster coredns-hanzodns service), preserving
verb, path, query, body, status codes and error bodies.
Isolation is bearer-relay: the head forwards the caller's OWN validated bearer
(cloud.CallerBearer) plus the server-validated X-Org-Id, substituting NO service
credential, so the DNS plane's own per-org authorization still holds and a caller
in org A can reach only org A's zones. Fail-closed: no validated principal => 403,
before any byte leaves cloud. It builds a fresh upstream request, so no inbound
header is blindly relayed; the upstream host comes only from env (no SSRF).
Decomplect the token resolution the identity boundary and this relay both need
into one callerToken helper (validatedPrincipal now delegates to it) and expose
CallerBearer for the relay; an opaque API key is never relayed as a bearer.
Assisted-by: Claude:claude-opus-4-8
The monochrome dashboard now serves at cd.hanzo.ai/ (root, base-href /) from the
cd-ui App CR (hanzoai/spa); cloud keeps ONLY the IAM-gated projection API at
/v1/deploy/api/*. Drops the go:embed FE + the deploy-ui-embed Dockerfile stage.
The FE is no longer /v1/-prefixed and no longer baked into the money binary.
The monochrome dashboard SPA now ships as the hanzoai/spa-based cd-ui App CR
served at cd.hanzo.ai/ (base-href /); cloud keeps ONLY the IAM-gated projection
API, moved from /v1/deploy/ui/api/* to /v1/deploy/api/* (same-origin with the
SPA). Removes the go:embed dashFS + static serve (dashStatic/serveDashIndex) +
webui/dist + the deploy-ui-embed Dockerfile COPY stage. Guard test updated to the
/v1/deploy/api/* routes (still 403 without SuperAdmin). RED invariant holds — the
API terminates in cloud behind IdentityMiddleware + the guard.
The platform build muscle (launchDirectBuild) clones an https git URL; the CLI
sent the bare `owner/name` positional verbatim, so `hanzo build luxfi/wallet`
failed server-side with "repo.url must use https". normalizeRepoURL expands a
bare owner/name to https://github.com/owner/name (the host for every
hanzoai/luxfi/zooai repo) and passes an explicit URL / scp-style remote through
untouched. Test: TestNormalizeRepoURL. Completes the IAM-login dogfood:
`hanzo build luxfi/wallet ... --sha <full> ...` now 202s + launches the build.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Publish routing is now org-scoped: a project publishes to
<slug>.<org>.<apex> (e.g. myapp.maxpower.hanzo.app) instead of the flat
global <slug>.<apex>. The slug namespace becomes per-org — two orgs can
own the same slug and their sites can never collide or shadow one another.
- deploy.go: siteHost(org,slug)=<slug>.<org> is the ONE bound/resolved key;
onPublish binds it; siteURL renders https://<slug>.<org>.<apex>.
- sites.go siteSlug: accept the two-label host <slug>.<org>.<apex>, validate
both labels (slug non-reserved), return <slug>.<org> as the resolve key so
bind and resolve agree. Org isolation is now STRUCTURAL in the hostname.
- store unchanged: site_hosts already keys on arbitrary full host strings
(custom-domain path proves exact full-host ResolveHost match).
containment.yml: add a required "zen streaming-fix floor" check that fails
any PR/push whose effective github.com/hanzoai/zen is below v1.4.1 (the
first release carrying the SSE body-close fix, commit 50328b8). A stale
branch that reverts go.mod's zen pin to v1.4.0 can no longer silently
re-break streaming (empty SSE completions) — the durable root-cause guard.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A plain `hanzo login` (IAM) now authorizes every PaaS control-plane op with no
separate --build-token / --platform-token. ONE identity, org+role scoped.
CLI (cli/cli.go): Env.buildToken() and Env.platformToken() fall back to the IAM
access token as the FINAL resort (precedence unchanged above it: flag > env >
credential-store service token > IAM login). So after `hanzo login` the CLI sends
the IAM JWT as the platform bearer. "No token" errors now point at `hanzo login`
and fire only when there is ALSO no IAM login. Tests: added
TestBuildTokenFallsBackToIAM / TestPlatformTokenFallsBackToIAM (precedence
preserved — a dedicated token still wins).
Platform (clients/platform/runner.go): /v1/runner (build-enqueue) — the one
control-plane endpoint that ignored identity — now accepts EITHER the shared
build-callback token (machine path: git-push, self-release, operator; constant-time,
unchanged) OR a validated IAM principal who is an admin (principal.IsSuperAdmin ||
principal.IsOrgAdmin over principal.Validated). Both bounded by the SAME
owned-registry allowlist, so identity never widens the image boundary. Release
self-publish stays machine-token-only. IAM builds are org-attributed to the caller's
VALIDATED org and refuse a foreign organizationId unless SuperAdmin. Reuses the ONE
identity verifier (SanitizeIdentity mints unforgeable X-User-* from the verified JWT)
— no parallel JWT crypto. deploy/apps were already IAM-authorized via the tenant()
boundary. Tests: 7 new IAM cases (admin launches, non-admin 403, forged-no-user 403,
disallowed-image 403, foreign-org 403, release 403) + corrected fail-closed 403.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The argo/gitops merge d3f60be (v1.801.77) resolved the go.mod conflict to its
stale second parent (zen v1.4.0), silently reverting the v1.4.1 bump landed at
v1.801.76. v1.4.0 still carries the serve() `defer resp.Body.Close()` race that
empties every streamed completion (200 with a 0-byte body) — which broke every
hanzo.app builder stream. Re-pin to zen v1.4.2 (the body-close fix + reasoning_content
passthrough regression pin) and thinking v0.1.1 (reasoning default). Streaming
now forwards every chunk, including delta.reasoning_content.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The deploy-ui-embed image is published; cloud's Dockerfile now COPYs its /dist
into clients/deploy/webui/dist (go:embed). release.yml builds a cloud image that
serves the REAL monochrome dashboard at /v1/deploy/ui instead of the fallback.
Mirrors the console-embed stage: FROM the prebuilt deploy-ui-embed image, COPY
/dist -> clients/deploy/webui/dist (go:embed source for /v1/deploy/ui). GATE: do
NOT merge until ghcr.io/hanzoai/deploy-ui-embed:latest is published (else the
build cannot pull the base). Until merged, the money binary serves the fallback.
LOW-2: TestDeployRoutesRequireAdmin asserts every /v1/deploy(/ui) route 403s
without X-User-IsAdmin and passes with it (health stays public) — guards against
a future unguarded-route refactor.
INFO-1: the unauthenticated /v1/deploy/health path reports booleans only; the raw
k8s error (apiserver/RBAC detail) is logged server-side, not returned.
LOW-1 (CSRF): verified no-op — the IAM session cookie is SameSite=Lax AND the
ambient cookie->JWT bridge is same-origin-gated (sessionBridgeSameOrigin), so a
cross-origin CSRF POST gets no identity and the deploy guard 403s.
Serves the full ArgoCD React UI fed a read-projection of operator App CRs shaped
as v1alpha1 Applications — no argocd api-server/repo-server/redis/stored CRD.
SuperAdmin-gated; IAM owns identity at the edge. Money binary builds green;
projection render tests pass. Real monochrome bundle is a CI artifact (make
deploy-ui / deploy-ui-embed image); fallback shell until that lands.
CI story for the dashboard bundle, mirroring make webui: DEPLOY_DIR=<hanzoai/deploy
rebrand/hanzo-monochrome> yarn build -> clients/deploy/webui/dist (go:embed). Only
the fallback index.html + .gitignore are tracked; the real 43MB bundle is
build-time-only. Money binary builds green with the real bundle embedded.
Serves the full ArgoCD React UI fed a READ-PROJECTION of operator App CRs shaped
as v1alpha1 Applications — NO argocd api-server, NO repo-server, NO redis, NO
stored Application/AppProject CRD. projection.go maps App CR -> Application +
resource-tree (reusing the native readers/engine health). dashboard.go
reimplements the UI's api-server subset (settings/userinfo/version/can-i +
applications list/get/resource-tree + sync/rollback->App-CR reconcile) + serves
the go:embed'd monochrome bundle with base-href rewrite. SuperAdmin-gated; argocd
auth disabled (IAM owns identity at the edge). Projection render tests green.
UI bundle is a CI artifact (committed fallback shell; make deploy-ui overwrites).
v1.824.2 unmasks the router-stats model ids (arm-N → real names like zen5-coder /
opus-4.8) for the PLATFORM scope when the caller is a super-admin of the own brand;
every other caller keeps the arm-N privacy masking. So the world.hanzo.ai admin view
(Routing Throughput / Enso arms) shows actual models instead of "Enso arm 2".
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
world's admin fleet (GET /v1/machines -> listMachines) sourced machines only
from Visor's registry (/v1/get-machines), so DigitalOcean droplets that were
provisioned but not (yet) in the registry never appeared -- DO nodes were
entirely missing from the fleet.
Source the managed-machine set as the deduped UNION of the registry AND
Visor's LIVE DO reseller list (GET /v1/machines -> ListComputeMachines ->
service.ListOrgMachines, the live Droplets.ListByTag(orgTag)). Dedup is by
provider id OR name; the registry entry wins a collision so its enrichment/
masking is preserved. One helper (managedMachines) now feeds listMachines,
listGPUs and the /v1/fleet board so all three agree on which machines exist,
not just how they normalize. BYO fold unchanged; only machines Visor actually
returns are surfaced (nothing fabricated).
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
buildMeteringClient ignored the documented METERING_TEST env, so the metering client
was ALWAYS live (c.test=false) — a staging/canary could not route debits to the
sandbox books, and the usage-cap smoke would have moved real money. Now METERING_TEST=true
sets Config.Test, so fin.RecordUsage writes the TEST finance books and the cap read
(org.TestMode via SQUARE_ENVIRONMENT=sandbox) sees the SAME test books. Unset in prod
= live, unchanged.
The spend cap read commerce's transaction store, which the co-resident cloud binary
leaves EMPTY (usage is recorded via fin.RecordUsage on the finance ledger) — so in
prod the cap summed 0 and never enforced, and the alert never fired. This wires the
cap onto the ledger prod actually writes, ORG-WIDE (the finance Entry carries no
scope; per-scope is a follow-up):
- sqlstore.SumByKindSince: additive read-only aggregate (kind + created_at index,
18-decimal TEXT folded in Go) — no Entry schema change.
- finance.SumUsageSince: the org's metered usage (cents) since a cutoff, deposits
excluded, sandbox books for a test org.
- finance.SetUsageHook: dependency-inverted post-debit seam (finance never imports
commerce) the cap alert fires through.
- apps/commerce.go: SetPeriodSpendReader(financePeriodSpend) so AuthorizeSpendCap
reads finance spend since the UTC month start, and SetUsageHook(fireCapAlert) so a
finance debit fires the org's spend-alerts on the same crossing.
Composes the commerce policy/CRUD/promo/admin/ancestor-fix (commerce
v1.49.1->v1.49.2 injection seam) — a targeted host re-wire, not a redo.
Completes the Enso per-tier gate: ai v1.824.1 already enforces min_tier at the
family pipe + auto-router; this bumps the co-resident commerce so /v1/billing/tier
returns the caller's REAL plan (was stubbed always-Free). Fail-open on uncertainty.
v1.824.1 boots StartRouterTrainer + StartRouterProbe from ai.Mount, so the flywheel
runs in the deployed (embedded-in-cloud) service, not just the standalone aid binary.
Both still self-gate on their env flags; universe sets ROUTER_TRAIN_ENABLED=1 to turn
training on. Also carries the retrain-timeline fix (retrains now count).
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
Five guards before any deletion: (i) refuse an empty desired set; (ii) dry-run
sizes the prune set + a count/ratio fuse (DEPLOY_ENGINE_PRUNE_MAX default 10,
_RATIO default 0.20) refuses a mass prune; (iii) WithPruneConfirmed gates prune
on the fuse passing; (iv) PVC + KMSSecret are excluded from prune entirely (data
anchors, irreversible); (v) parseManifestDir walks recursively so a nested
manifest is never silently dropped (which prune would read as a deletion).
prune stays off by default (DEPLOY_ENGINE_PRUNE).
Drops the filesystem replace => ../deploy/gitops-engine. The fork's engine module
was renamed to its real repo path (github.com/hanzoai/deploy/gitops-engine, tag
gitops-engine/v0.7.2) so cloud requires it as a normal pinned version — CI builds
the money binary with NO sibling checkout, NO argoproj alias. tidy + scoped build
green over SSH.
Every merge to main builds + smoke-tests + tags a proven image, but nothing
recorded that tag as the desired state Hanzo CD deploys, so api.hanzo.ai sat on
a stale pin (v1.801.71) while proven images (…72-…75) never rolled. The old
image-update.yml deploy hub was deleted in the Hanzo CD cutover; a direct CR
patch is reverted by ArgoCD selfHeal.
Add a promote job that, after the tag receipt, bumps spec.image.tag in
hanzoai/universe crs/cloud.yaml and commits deploy(cloud): <tag> — the SAME
yq-bump the hanzoai/ci reusable does for every other service. The universe-crs
ArgoCD Application (automated sync + selfHeal) then reconciles it to the cluster.
No hand-dispatch, no hand-edit.
Fixes empty streaming completions for zen5* (hanzo.app builder P0):
- zen v1.4.1 stops serve() from closing the upstream body before fiber's lazy
SendStreamWriter drains it (every SSE completion was truncated to a 200 + empty body).
- thinking v0.1.1 makes glm-5.2/deepseek Off send reasoning_effort:none, so zen5-coder
streams the answer immediately instead of a long silent content:null reasoning preamble.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds the /v1/admin control plane admin.hanzo.ai drives, twinning /v1/admin/flags:
- /v1/admin/promos (GET/PUT, core.Guard SuperAdmin) → commerce /v1/platform/promo:
configure the admin-controlled plan promo (percentOff/start/end/plans/active).
- /v1/admin/spend-caps (GET/POST/PATCH/DELETE, core.GuardScoped) → commerce
/v1/billing/spend-alerts with X-Org-Id: oversee/override ANY org’s usage caps
(SuperAdmin via ?org=; a scoped admin hard-pinned to their own — the escalation
line). Reuses the customer’s OWN spend-alert rows, no parallel model.
commerce.Forward is the ONE service-token seam these ride, relaying commerce’s own
status so a 400/403/404 surfaces honestly instead of masking as success.
Fixes clients/metering/middleware.go defaultOnDenied: a FUNDED caller over a
per-scope spend cap now maps to a DISTINCT 402 spend_cap_exceeded (errors.Is), not
the 503 it fell through to — parity with the zip-native denyVerdict, so any product
on the net/http middleware surfaces the same honest verdict.
Brings the ACTIVE Hanzo KMS custody plane (api.hanzo.ai/v1/kms/*, served by the
cloud-embedded luxfi/kms per HIP-0106) to the latest v1.x. keys (v1.4.1) and
crypto (v1.20.2) already latest; luxfi/mpc stays out of the graph (threshold
signing is a wire-coupled external daemon, not a linked module). v1.12.4 verifies
against the public sumdb; clients/kms + cmd/cloud build green.
toMachineView's memGB fallback read the integer before "gb" out of any
size slug. On a DO GPU droplet the slug's gb is VRAM (gpu-h100x8-640gb ->
640 GB VRAM), not system RAM, so a GPU node missing its upstream memSize
would render 640 GB of "system memory" -- a misleading number.
Guard the fallback with the GPU check already needed for v.GPU: reuse the
single gpuSpecOf(slug) call (spec, isGpu) and apply the slug's gb figure
only when !isGpu. Real m.MemSize still takes precedence for every provider,
GPU nodes included, so a GPU machine that reports its true RAM is
unaffected -- only the VRAM-as-RAM fallback is suppressed.
Table tests: a gpu-h100x8-640gb slug with empty MemSize yields Mem=="" (not
"640 GB") while still resolving GPU=="H100", and the same slug with a real
MemSize=="1920gb" reports "1920 GB".
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
The fleet view (world.hanzo.ai cloud variant) renders machines from
cloud's /v1/machines -> listMachines -> toMachineView. Two honest-data
gaps left system RAM and DigitalOcean vCPU counts blank:
1. System memory was never surfaced: machineView had no memory field and
toMachineView never read the upstream memSize, so every provider's RAM
column rendered empty.
2. DO vCPU was dropped: toMachineView filled vcpu only when CpuSize parsed
as a bare integer, but DigitalOcean reports size SLUGS (s-4vcpu-8gb),
so strconv.Atoi failed and vCPU showed nothing.
Fix:
- Add MemSize to visorMachine (upstream already sends it; it was simply
unmapped) and Mem to machineView.
- parseSizeSlug pulls the integer before "vcpu" and the integer before
"gb" out of a size slug (s-4vcpu-8gb -> 4,8; g-8vcpu-32gb -> 8,32).
- normalizeMem renders "N GB" only for trustworthy inputs (explicit
gb/gib, explicit mb converted with rounding, a bare integer as MB when
>=1024 else GB) and returns "" for anything ambiguous -- never a
fabricated number.
- toMachineView keeps the Atoi(CpuSize) path and falls back to the slug's
vcpu; sets Mem from normalizeMem(MemSize) and falls back to the slug's
gb figure. The GPU-spec logic is unchanged.
Table tests cover the slug parser, the mem rounding, and the mapper
precedence (explicit values win, honest omission when neither yields one).
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
collectOutputs read only the 'images' key, but SaveGLB publishes under '3d', so a
generated .glb never mirrored to the org library. Gather every saver's outputs
(images + 3d), so the studio 3D lane's mesh lands like an image or video render.
The cloudflare provider now offers a browser OAuth path in addition to the
shipped apikey path. /connect dispatches by request: a "token" key in the body
seals via apikey (verify-before-store, unchanged); its absence starts the
Authorization Code flow (confidential client, client_secret, no PKCE — the
framework's OAuth pattern) and the exchanged access token is sealed to the SAME
KMS coordinate (/orgs/{org}/integrations/cloudflare/api_token), so the DNS
provider layer is auth-method-agnostic.
Framework: connect dispatch is now capability-based (Verify and/or Authorize)
rather than Kind-only; Mount validates RedirectPath for any OAuth-capable
provider; bodyHasCredential picks the path. The OAuth leg gates on its own app
creds (Creds().ClientID) so a missing Cloudflare OAuth app degrades to an honest
503 without breaking the always-available apikey path.
Requires a registered Cloudflare OAuth app: CLOUDFLARE_OAUTH_CLIENT_ID/SECRET in
env, redirect https://api.hanzo.ai/v1/integrations/cloudflare/callback.
No Go change — this rebuild re-resolves the freshly-republished console-embed:latest
(console main bd8a81651) into the go:embed console served at console.hanzo.ai, and
ships builtin.go's auto '/v1 route → MCP tool' plane at /v1/tools/mcp (already in main,
newer than the deployed v1.801.69).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The value routes registered the optional-greedy wildcard `/secrets/*`, which
fiber also matches with an empty tail — so the bare `GET .../secrets` list path
was answered by getSecret (400 "secret name is required") and listSecrets was
unreachable. Switch the getSecret/deleteSecret value routes to the required-
greedy `+` (one-or-more), so `/secrets` falls through to the exact list route
while `/secrets/<path>/<name>` still reads/deletes. reqWildcard reads the `+`
param.
Regression test list_route_test.go asserts the bare list path returns 200 with
a secrets array (was 400) and that value reads still work.
Register Cloudflare as an apikey-kind connector on the /v1/integrations plane.
A customer-supplied scoped API token is verified live against Cloudflare's
GET /user/tokens/verify (must be status:active) before it is sealed into the
org's KMS namespace (/orgs/{org}/integrations/cloudflare/api_token); the
connection row holds only non-secret account metadata. connect/verify/disconnect
are org-admin gated from the validated principal (principal.IsOrgAdmin).
Extends the connector framework with the apikey credential seam shared by future
customer-credential providers: Provider.Kind/AdminOnly/Verify, VerifyInput,
connectByCredential (verify-before-seal, fail-closed), and the verify route.
Serves POST /v1/integrations/cloudflare/{connect,verify,disconnect} and
GET /v1/integrations.
The cloud already trusts hanzo-admin-guard (the admin surface) but not admin-console
(the admin console's own OIDC client), so a SuperAdmin token minted via admin-console
was rejected on /v1/admin with 'invalid audience' — forcing an awkward hanzo-admin-guard
detour. Add admin-console so the admin console's tokens work directly, matching
GATEWAY_ALLOWED_AUDIENCES which already lists it.
Co-authored-by: zeekay <z@hanzo.ai>
Plane map with honest tiers in LLM.md: /v1/connectors custody (in flight), /v1/channels transport (planned, branch reserved), shipped planes named by package. Spec home HIP-0129; roadmap P1-P15 lives there.
"gatewaypolicy" is a compound (gateway+policy) and read as a second gateway
package. It is the ONE Gateway concern with clients/gateway — the per-org edge
policy STORE (OrgRPM ceiling + CORS + cache) that the /v1/gateway/config plane and
the package-cloud edge middleware both read.
They are two packages only to break a Go import cycle: clients/gateway imports root
cloud (cloud.Deps), and middleware_edge.go IS package cloud — so the store must be a
LEAF both can import. A flat merge cycles. Fix per the no-compound law: nest the leaf
UNDER gateway as clients/gateway/edge (edge.Policy/Store/New). One gateway namespace;
/v1/gateway/config surface unchanged; edge-middleware logic unchanged.
NOT redundant with the external hanzoai/gateway (KrakenD): that does coarse per-route
edge rate-limit + auth at ingress; this is per-AUTHENTICATED-org RPM (needs the decoded
token org), an app-level ceiling the edge proxy cannot compute. Different layer.
Pure rename (49/49, no logic change); full cmd/cloud binary links; gateway + gateway/edge
tests pass; gofmt/vet clean.
waitlist_* / public_signup / gateway_* are runtime flags; strip their Env: fallbacks so
they resolve from the /v1/flags DB engine → Default (single source of truth, flipped live,
no redeploy). Boot-time ReadOnly rows (subsystem_*, network_id_*) keep Env — that IS their
boot mechanism. Nothing read these env vars outside the flags engine (verified).
Co-authored-by: zeekay <z@hanzo.ai>
Wire the load-bearing trigger so a coding run can be dispatched to a chosen
linked machine end-to-end. Extend the Slack coding grammar with an optional
routing prefix — code: <repo> on <machine> <task> — resolve <machine> org-scoped
(id or friendly label) and set Req.TargetID. An unknown or foreign machine is an
honest error, never a silent local fallback; an untargeted request is byte-
unchanged (repo <task>).
- agents.ResolveTarget: the ONE org-scoped id-or-label resolver, fail-closed, so a
trigger surface turns 'on evo' into a target id without leaking another tenant's
inventory.
- slack_coding: parseCoding yields (repo, target, task); a routed run skips the KMS
agent-credential fetch (the machine authenticates with its own credential); the
ack + result card report a routed run as queued-on-<machine>, followed live in
mission-control, not a premature branch-pushed verdict.
Tests: ResolveTarget id/label precedence + cross-org not-found + unmounted fail-
closed; parseCoding on-prefix grammar (routing only when 'on' is the token after
the repo; 'on-call'/'only' untouched); routed result card is queued not done.
cloud transitively pulled github.com/WqyJh/{go-cosyvoice,go-openai-realtime}
through ai v1.822.2 (the go-openai-fork release, which predated ai's TTS
switch to the hanzo-owned forks). ai v1.822.3 wires ai/tts onto
github.com/hanzoai/go-cosyvoice + go-openai-realtime, so tidy drops both WqyJh
edges from cloud's graph. cloud now pulls ZERO third-party OpenAI-lineage:
WqyJh 0, sashabaranov 0, ClickHouse 0. Single datastore sql registrant
(hanzo-ds/go). Builds clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
admin.hanzo.ai's SaaS-metrics/Invoices/Subscriptions pages were placeholders
awaiting cross-org /v1/admin/* endpoints. Add them as super-admin (core.Guard)
fleet-aggregate readers over the existing commerce billing engine:
GET /v1/admin/metrics — SaaS god-view: MRR/ARR/net-new/churn/active-subs/
paying-customers/plan-mix/top-customers/recent.
Single S2S proxy — commerce /v1/metrics/saas is
already a cross-org aggregate (same gate finance
Costs uses: RequirePlatformAdmin→IsServiceToken).
GET /v1/admin/invoices — cross-org invoice list; fan core.ListOrgs out
GET /v1/admin/subscriptions — cross-org subscription list; per-tenant reads
merged (identical to revenue.go's fan-out).
Honest degradation: a failed per-org read contributes no rows, never fabricated.
go build ./clients/admin/... green, gofmt clean. Pairs with admin operator UI
(feat/admin-billing-fleet-ui).
cloud/clients{,/agent} used sashabaranov/go-openai only via a 'replace =>
hanzoai/go-openai' — which does not propagate to cloud's own consumers, so
gateway/iam/etc. each had to copy it. Now the fork declares its own module
path (github.com/hanzoai/go-openai v1.41.0): require it directly. Bumps the
lockstep fork adopters — ai v1.822.2, agent v0.1.3 — so the hz.Mount Completer
boundary shares ONE openai type set (was a hanzoai-vs-sashabaranov type
mismatch). No replace anywhere; upstream sashabaranov remains only as an
indirect dep of the go-cosyvoice TTS chain (owned next). cmd/cloud keeps its
single datastore sql registrant (hanzo-ds/go). Builds clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Embeds the fleet KMS into cloud by re-sealing the ~125 KMSSecret-referenced
secrets from the legacy standalone (unsealed at rest) into cloud's embedded
/v1/kms (AES-256-GCM sealed per secret). Driven by the KMSSecret CRs — the
authoritative (org,path,env,key) manifest — not a raw store copy.
- inventory: pure CR parse/validate/dedup; handles explicit keys, folder-sync
(empty keys[]), the env-default divergence (cloud refuses empty env), and
malformed CRs (fail-loud, never silently dropped). Reuses cloud/clients/kms
ValidSegment/ValidSubpath so a coordinate the store would reject is never built.
- reseal: per-CR org-bound auth (owner==projectSlug), GET standalone -> POST cloud
(cloud seals). Idempotent upserts, re-runnable. Plaintext transits memory only,
wiped after write; results carry coordinates + status, never values.
- verify: read-only SHA-256 hash-compare standalone-vs-cloud + org-isolation matrix
(cross-org 403, no-principal 403) against the real cloud guard.
- preflight: cloud /v1/kms reachability + JWT-validation probes + offline G1 delta.
- runbook: the ordered, rollback-safe cutover (standalone stays read-only).
Tests: round-trip against cloud's REAL embedded /v1/kms in-process (real seal,
real guard) via a zip app.Fiber().Test transport adapter; seal-proof (no plaintext
on disk); wrong-org refusal; folder-sync via LIST; hash-mismatch detection;
isolation matrix. go test ./cmd/kmsreseal green.
featuregate READ like a synonym for flags — the source of the "isn't this the
same thing?" confusion. It is not: flags is the Policy decide-ENGINE (/v1/flags);
this package is the request-ADMISSION gate that composes it one-way (host→service
registry + waitlist.<svc> mode read + Enforce middleware + IAM approval check).
Renamed to `admission` — the precise systems term for policy-gating requests
(k8s-style admission control) — which also dodges the gate/gateway/gatewaypolicy
name cluster. flags stays THE engine; admission is a thin one-way consumer.
Also kills the /v1/featuregate/mode compat alias entirely (route + Enforce exempt
entry + test): one route, /v1/flags/waitlist. No shim, no adaptor, no backwards
compat — per the one-and-only-one-way law.
flags engine surface unchanged. Build/vet green; admission tests + apps frozen-Wire
order test pass (admission holds featuregate's slot). Deeper Policy collapse
(authz/entitlements/gatewaypolicy → one engine) is a separate staged HIP-0127 pass.
When a coding run carries a targetId (a registered /v1/agents/targets
machine), enqueue it as a durable task addressed to that target on the ONE
embedded tasks engine instead of running it in the cloud sandbox. No target
keeps the local sandbox path byte-unchanged.
- mailbox: an in-process rendezvous between the durable RoutedRunWorkflow and
the external machine that claims a run over HTTP; tenant + machine isolation
is a property of the (org,target) key, not a check a caller can skip.
- routing: per-target claim key (a capability, stored only as a SHA-256 hash,
constant-time verified) is the machine identity; a fail-closed liveness gate
(online + a live runner) is the dispatch admission.
- claim/report HTTP surface (org bearer + X-Target-Key) lets a machine claim
and complete only runs addressed to it.
- coding.Dispatcher gains a routed branch: open the session on the target,
enqueue the durable RoutedRunWorkflow (no secret in the payload — the machine
authenticates with its own credential), return queued; fail closed on an
unavailable target or a failed enqueue, never fall back to local.
Tests: dispatch-to-target, no-target-local-unchanged, dead-target-fail-closed,
cross-machine/cross-org claim denied, mailbox isolation + claim race.
Swap the S3 client in the 7 direct importers from github.com/minio/minio-go/v7
to github.com/hanzoai/s3-go (package minio; a minio-go v7.0.98 fork). Drop-in:
same package name and same New/Client/Options, {Get,Put,List,MakeBucket,
RemoveObject,RemoveObjects}Options, ObjectInfo/Object/ErrorResponse surface and
credentials.NewStaticV4; conditional-CAS (SetMatchETag/SetMatchETagExcept) and
presign paths unchanged.
minio-go leaves the direct requires and stays indirect (luxfi/zapdb via
clients/kms). Run go mod tidy after the s3-go v1.0.0 tag is published to
populate go.sum.
The supervisor recycles at queue-idle, but a freshly CLAIMED job is invisible
to the engine queue until its graph is submitted — so recycles fired over the
claim-to-submit window and staging failed on a dead engine, consuming the job
(observed twice in prod, seconds apart). Two invariants close it: a staging
latch the supervisor honors before recycling, and waitEngine() so a job
claimed while a recycle is already mid-flight waits out the restart instead
of dying on connection-refused.
The account-usage plane (7456318) wrongly opened a SECOND usage surface inside
clients/link (/v1/links/usage). Move it into clients/usage so usage owns ALL
usage and link owns links and nothing usage — one surface, orthogonal, one window.
Moves (package link -> usage): sample.go (the Sample value + Sanitize), datastore.go
(the hanzo.account_usage warehouse series + reads, now behind a `warehouse` type that
holds only the DDL latch over aiobject's shared datastore — no handle, so usage keeps
NO Shutdown), and the record/samples handlers (account.go). Reconciled with the usage
subsystem: cloudUsageTable -> the existing llmTable, dsTime -> the existing tsLiteral,
duplicate aString -> dsString.
Route table (was /v1/links/usage*):
POST /v1/usage record account-usage samples (the collector)
GET /v1/usage/samples one provider account's own lane dash (time series)
GET /v1/usage/summary THE one summary — merged (see below)
GET /v1/usage/analytics{,/access} unchanged
Summary collision resolved by MERGE, not two endpoints: the account-usage global view
folds into the existing /v1/usage/summary as a labelled `accounts` block beside spend +
LLM, over ONE window (aiobject.ResolveCloudUsageWindow drives both). Nothing dropped —
the caller's own linked-account rows AND the org Hanzo-routed rows both ride the one
summary, each side reporting its own availability, never summed.
Decomplected the Link-refresh: reportUsage braided a warehouse write with a Link
upsert, and since POST /v1/links already sets an account's usage snapshot, the sample
-> snapshot path was a SECOND way to do that. record now records usage only; the link
registry stays link's own concern. Drops the 3 Link-registry tests (they exercised
/v1/links, unreachable in a usage-only mount) and the Link half of 2 more; the warehouse
+ value coverage moves intact. No back-compat alias (the route was hours old).
Wire guard unchanged: link keeps its Shutdown (SQLite store), usage keeps none.
console.hanzo.ai is served one-binary off cloud, so its /v1/iam/* calls hit the
iam_edge — which required a validated org for EVERY route. That 401'd
'sign in to continue' on the sign-in routes themselves (get-app-login, login,
oauth token exchange), a chicken-and-egg that bricked console login (the
'unknown iam route' / 'sign in to continue' users saw). Forward the
unauthenticated-by-design sign-in surface (login-page config, credential submit,
signin/signup, captcha/verification aids, the OAuth token endpoint, OIDC
discovery) straight to IAM BEFORE the org gate. Tenant CRUD + org metadata stay
fully gated — no tenant-data route is opened. Test: TestIamEdgePublic.
v0.15.4 closes the red-team CRITICAL in the federation broker: authorize
Application.Organization on write + reserved-org guard in federation
link/provision (was: social login could mint a SuperAdmin / take over a
cross-tenant account) + SSRF IP filter. Required before the hanzo.id social
cutover. Build pipeline healthy (consensus v1.36.3).
Assisted-by: Claude Opus 4.8 <noreply@anthropic.com>
A read-only subsystem_iam2_active switch on the platform panel, mirroring
subsystem_iam_active, makes the clean-room iam2 selection visible in the /v1/flags
cockpit. The selector stays ONE thing — CLOUD_IAM_IMPL=iam2 at boot, applied on
the next reconcile — this switch reflects it, it does not add a second control.
Its description names the gate: the IAM cutover parity suite
(universe e2e/50-iam-cutover-parity) must be green against the iam2 shadow before
the canary is flipped.
The account-usage plane over clients/link: a Sample value (one metering lane of
one provider account at one instant), a ReplacingMergeTree warehouse projection
(hanzo.account_usage + a dedup-preserving daily rollup MV) read back with explicit
read-time argMax dedup, and the /v1/links/usage surface — report samples, a
per-provider dash, and a global summary that sets a user's own linked-account plan
usage beside the org's Hanzo-routed cost of record, every row labelled by
source/scope/confidence and never summed together.
A windowless sample (a valid window class with no meter-reported duration or
reset) keys its class's nominal bucket, never the zero instant: every ranged read
filters window_start into [from,to) and the TTL drops epoch rows on arrival, so a
zero-keyed row would be written-but-never-read and would silently drop out of the
summary. Re-polls of a windowless counter collapse onto that one nominal instance
(ReplacingMergeTree by ts), so it is one row per lane, never summed across polls —
reconciling the two window-instance tests the rescued WIP left in contradiction.
The header was captured from client input and re-injected verbatim for any
validated principal. That was defensible while it was a mere attribution hint
no debit ever read — the comment said as much. It is not one anymore:
ai/object.Payer now resolves the PAYING account from it, so forwarding the
client's copy would let a caller name its own payer, which is the whole thing
the claim exists to prevent. A signup-org member could have sent
`X-Billing-Account-Id: org:hanzo` and pointed their spend at the shared pool.
It is now minted from the validated `billing_account` claim
(idClaims.mintedBillingAccount), mirroring iamauth.Claims.MintedBillingAccount
byte-for-byte, so the in-binary path binds what the gateway would and both
resolve one payer. The raw client copy is deleted on ingress and not restored.
The console read and the top-up now hand Payer that same claim, so the balance a
member SEES, the account a top-up FUNDS, and the account the ai gate DEBITS are
one wallet. Feeding Payer a different credential per call site is the modern
shape of the old org-vs-"org/user" split: a funded balance the gate refuses.
Tests drive real signed tokens through the boundary: the claim reaches the
header for person/org/project, a forged copy never survives (even on a token
that carries no claim, where a restored copy would be the only value present),
and an anonymous caller carries no payer at all.
flags is now the PURE (Principal,context)->verdict engine: Register/Bool/Int/
String/Board/SetPlatformSwitch/Defs + /v1/flags/* + native evaluator + the
platform-switch seed. ZERO host->service / ModeForHost / waitlist.<svc> /
mode-route knowledge.
The complete launch waitlist-gate feature moves to clients/featuregate, which
COMPOSES flags one-way (flags.Bool/Register/SetPlatformSwitch/Def/Defs; flags
never imports featuregate):
- flags/waitlist_store.go -> featuregate/registry.go (host->service map)
- flags/waitlist.go -> featuregate/waitlist.go (mode decide + admin funcs +
seed + waitlist.<svc> Def registration + Mount/Shutdown + the mode route)
- flags/waitlist_store_test.go -> featuregate/registry_test.go
- the registry OrgStore handle (was flags.Client.registry) is now featuregate
package state, opened in featuregate.Mount, closed in Shutdown
- Enforce default gate is now the LOCAL WaitlistModeForHost
- /v1/flags/waitlist AND /v1/featuregate/mode compat alias served by
featuregate (route name unchanged)
- apps.Wire re-adds featuregate after admin (after flags); wire_test frozen row
- admin/services.go swaps the flags import to featuregate for the board funcs
The namespace collapse (75d6f36) renamed the live waitlist-mode read to
/v1/flags/waitlist and made /v1/featuregate/mode a 404 — correct per the one-namespace
Policy primitive, but a BREAKING change to a public route whose external callers cannot
be fully enumerated from the monorepo (a deployed frontend could still call the old
path). Per the hard "never goes down for any customer" constraint, ship the collapse
WITHOUT the break: /v1/flags/waitlist is canonical; /v1/featuregate/mode is a temporary
alias to the same handler; both exempt from the Enforce gate.
Delete the alias (this route + its exempt entry in featuregate/middleware.go) once every
caller is confirmed on /v1/flags/waitlist — a one-line follow-up, gated on the owner.
Verified: gofmt clean, go build/vet green, exempt-path test asserts BOTH routes ungated.
services.hanzo.ai is dead (0 Service CRs cluster-wide; the fleet is 100% App).
clients/paas, clients/deploy, and clients/platform drop the two-kind read shim
and read apps.hanzo.ai only. The paas deploy endpoint (and release seam) now
always refuse a git-declared App with 409, naming the universe git path to
commit the tag to.
The services.hanzo.ai kind is dead: zero Service CRs exist cluster-wide and
the whole fleet is apps.hanzo.ai (kind App). These three cluster-facing planes
carried a two-kind read shim (App first, Service fallback) that is no longer
reachable, so strip it and read one kind — App.
- clients/paas: drop servicesGVR + crGVRs(); listApps/getApp/observeFleet read
appsGVR directly (no cross-kind dedup). The deploy endpoint now always refuses
(409): every App CR in the platform namespaces is git-declared and reconciled
by Hanzo CD with selfHeal, so a patch here is reverted — the response names the
git path to commit the tag to. releaseService refuses on the same grounds.
- clients/deploy: drop servicesCRGVR + appCRGVRs() and the "hanzo.ai/Service"
registry entry; health/getAppCR/listAppCRs read appsCRGVR directly. coreSvcGVR
(the core/v1 Service child object) is unchanged.
- clients/platform: drop servicesGVR + crGVRs(); resolveCR/getCR/deleteService
read and delete appsGVR only. Tenant apps are still written and patched as App
CRs in tenant-<org>.
Tests updated to the one-kind reality. Builds/vets/gofmt clean; go.mod untouched.
The guard's public waitlist-mode read now lives under /v1/flags (the flags engine
owns it) — there is NO /v1/featuregate HTTP endpoint. The route, the Enforce
exempt prefix, and the doc/prose comments move; the featuregate Go PACKAGE (native
Enforce middleware) is NOT renamed, and the /v1/admin/services board is unchanged.
- clients/flags/routes.go GET /v1/featuregate/mode -> GET /v1/flags/waitlist
- clients/flags/waitlist.go doc comments repointed
- clients/featuregate/middleware.go defaultExemptPrefixes /v1/featuregate/ -> /v1/flags/waitlist
- clients/featuregate/middleware_test.go exempt-path assertion updated
- apps/apps.go stale prose comment repointed
Verified green: go build ./clients/flags/... ./clients/featuregate/... ./apps/...,
go vet, and CGO_ENABLED=0 go test ./clients/featuregate/...
Three coordinate-hygiene fixes so the pipeline resolves deterministically:
- ai v1.820.0 -> v1.821.1. v1.820.0 pinned iam at an orphaned pseudo-version
(a commit rebased out of existence); v1.821.1 pins the real iam tag v1.31.28.
- iam -> v1.31.28, the real published tag; the pseudo-version and its replace
are gone.
- luxfi go.sum re-recorded from the immutable sum.golang.org via go mod tidy,
so consensus/vm can no longer carry the hashes a force-moved git tag served.
Nothing but coordinates changed; ai v1.821.1 is v1.820.0's tree with one dep line
repinned, so the compiled result is identical to the shipped v1.801.49. Real
public semver only: no pseudo-versions, no replaces, no force-moved tags.
The v1.801.50 tag collision: a run pushed :v1.801.50 then was cancelled after
imagetools-create but before its git tag (orphaned container tag). The Tag steps
container-tag floor (cont_max) was fail-OPEN — `gh api ... 2>/dev/null || true`
yields "" on any API error — so a later run did NOT see :v1.801.50, recomputed
the same number, and REASSIGNED :v1.801.50 to a different image: an ambiguous
mutable prod tag (silent flip on any fresh-node reschedule).
Fail-CLOSED: if the container-tag lookup ERRORS (vs legitimately empty), retry the
whole attempt instead of proceeding on a git-only floor that cannot see the orphan.
A version that already has a pushed image is now never reused.
NOT reordering git-tag before imagetools-create (the other candidate fix): that
reintroduces the phantom "tag exists, image does not" this workflow was built to
prevent. Pairs with the crane-mirror timeout (ed4d372) that stops the hang→cancel
which orphans tags in the first place. Compute-step cont_max left as-is (hint only).
[skip ci]
Recycling on each completed render killed long renders mid-sample when short
jobs shared the engine (observed: every direct render died within ~6 minutes
while probe jobs cycled). The recycle now defers until the queue is empty.
Health: an engine that answers /queue with work in it is alive however slowly
it answers /system_stats; restarts require three consecutive silent probes
with an idle or unreadable queue.
The 'Mirror credential' step fail-safed only on a missing KMS token, not on the
docker-login to registry.hanzo.ai itself. A transient 502 from the mirror registry
(ingress blip; the registry was healthy 6m before and after) killed the whole
serialized release — no image, no tag — even though ghcr (the PRIMARY) was fine.
Both login paths now skip the mirror (MIRROR_OK unset) on failure and continue.
Complements ed4d372 (the crane-copy timeout): the mirror is now best-effort end to end.
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
An unbounded `crane copy` to registry.hanzo.ai can HANG (not just fail) — the
best-effort mirror once livelocked the Tag step and held the entire serialized
release lane (concurrency: release-cloud, cancel-in-progress:false), so no queued
release could run. A best-effort mirror must never be able to block the git-tag
receipt that follows it. `timeout 120` makes it truly best-effort.
[skip ci]
v0.15.0 adds the OIDC/OAuth2 social-federation broker (Google/GitHub);
v0.15.1 mints signing keys for keyless reserved-org certs so the embedded
iam2 publishes a JWKS and can sign tokens (shadow-canary finding). Carries
the full parity + RFC surface into the cloud image for the hanzo.id cutover.
Assisted-by: Claude Opus 4.8 <noreply@anthropic.com>
luxfi/consensus v1.36.2 was force-repushed with different go.mod content, so
cloud's committed go.sum no longer matches and 'go mod download' aborts with a
SECURITY ERROR — breaking EVERY release. Same recurring luxfi force-move pattern
as c93ddf9 (keys). Bump to latest stable v1.36.9; clients/controlplane (only
importer) compiles clean, go mod verify passes.
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
* render: poll window matches the dispatch cap; engine recycles after each render
The 10m local history poll undercut the 4h startToCloseTimeout the dispatcher
grants — live renders (observed 8-70m) were marked failed while still sampling;
only the mirror later delivered them. renderWindow now matches the cap.
The engine leaks ~58GB per render. The handler signals a recycle after each
COMPLETED render (never on timeout — the engine may still be sampling and the
mirror rescues late finishes); the supervisor restarts on the signal.
* mirror: skip hidden files — AppleDouble forks pass the extension check
._foo.png is a mac resource fork, not a render; 700+ of them poisoned a
library within an hour of the mirror going live.
* deps: luxfi/consensus v1.36.2 -> v1.36.3 — the v1.36.2 tag was re-pushed
Cold builds fail sumdb verification against the moved tag (downloaded
eKzasq4O... vs sealed IbeWQF1w...). v1.36.3 is the immutable successor;
never re-tag a published version.
Ships to prod: router.enabled=true (model=auto routes for every org by default),
per-request RoutingEvent recording for auto AND explicit models (up/down feedback
works on all models), per-org + global fit-gate-deploy-publish training, and the
self-scoped routing-data export/delete (data ownership). Pairs with the universe
CR ROUTER_ENDPOINT removal (heuristic 300ns is the live path).
Claude-Session: https://claude.ai/code/session_018PmFAHZvbBSTsuWyebwMra
referrals/affiliates/authors each carried a byte-identical commerce.go (their own
doc-comments said so): the same commerce interface, httpCommerce, newCommerceClient,
deposit(), spendCents(), errUnconfigured — the S2S COMMERCE_SERVICE_TOKEN money-in
path (POST /v1/billing/deposit) + usage-rollup, triplicated.
Extract ONE clients/payout (attributed-credit -> commerce via commerceinproc): the
exported Commerce/Client/NewClient/ErrUnconfigured. Each program keeps a THIN
adapter — its own lowercase commerce interface + a commerceSeam that delegates to
payout.Client — so the program store/handler code AND their fakeCommerce test doubles
are untouched, and each program still names its own grant tag (grant:referral /
grant:affiliate / grant:author). ~330 duplicated lines collapse to one binding.
Zero behaviour change: identical HTTP contract, headers (X-Org-Id, Bearer), body,
fail-soft (ErrUnconfigured on deposit / 0 on spend when unwired), and errors.Is
sentinel. Adds payout unit tests (httptest) that give the extracted HTTP path REAL
coverage the fakes never did — ok clients/payout 0.010s.
gojabase was the read-WRITE-Base sibling of goja: it wrapped a goja.Host and
added per-tenant Base/SQLite persistence, but duplicated the Host/Config/Request/
Response/New surface. Fold it into the ONE goja package as the Base-binding
CONSTRUCTOR — the persistence layer is now opted into via NewBase (vs New for a
read-only catalog bundle):
goja.New / goja.Host / goja.Config / goja.Request read-only engine (plans/pricing)
goja.NewBase / goja.BaseHost / goja.BaseConfig / goja.BaseRequest + per-tenant Base
Moves gojabase.go -> clients/goja/base.go (renamed types, no goja. self-import),
store.go -> basestore.go, and both test files, all into package goja (zero
identifier collisions, coverage preserved). Repoints every importer —
dataroom/captable/sign (RW) to goja.Base*; plan/pricing already used goja and are
unchanged; base uses goja.TenantSegment. clients/gojabase deleted.
Behaviour is byte-identical: the engine, the per-request transaction commit-on-
<400, the injective TenantSegment, and the __db/__blob/__newId/__now host globals
are unchanged; only the package + exported names moved. No routes (both are
libraries). The gojabase[...] error prefix is kept as the RW-layer diagnostic label.
clients/connectorruntime mounts exactly ONE route —
POST /v1/automations/connectors/:id/run — the in-process goja runner paired with
automations own GET /v1/automations/connectors catalogue. It was a separate Wire
entry solely for that route. Fold connectorruntime.Mount in as a terminal
sub-mount of automations.Mount and drop its Wire entry + import -> ONE
automations subsystem.
The route is DISTINCT from every automations route and automations mounts no
/v1/automations/* wildcard, so there is no shadow; the runner still resolves the
shared engine lazily. clients/connectorruntime stays a focused package
(composition); its internal bundlecmd tool is untouched. Frozen wire row
removed.
clients/cron mounts NO routes — its Mount only launches a background starter
that registers durable schedules on the SAME shared engine (cloud.EmbeddedTasks)
that clients/tasks fronts. It was a separate Wire entry purely to get its
goroutine launched. Fold it in as a terminal sub-mount of tasks.Mount and drop
the cron Wire entry + import -> ONE tasks subsystem.
clients/cron stays a focused package (composition, not code-dumping): tasks
imports and invokes it. No routes change (cron never had any); the scheduler
still waits for the post-MountAll engine, so timing is unchanged. Frozen wire
row removed.
clients/plan.Mount was wired under the name "plans" while its package, and now
its generated standalone cmd, are "plan" — one subsystem, two names. Normalize
the Wire enable id (and cmd/plans -> cmd/plan, ServeSingle arg) to "plan".
Product routes are unchanged: the subsystem still serves /v1/plans/* (plural),
including its OwnsHealth /v1/plans/health probe — only the enable id / binary
name changes. No route drop; mount-all default still enables it (empty Enable =
all on). Updated the frozen wire row and the two cmd/cloud enable-id references;
TestMountAllAndServeHealth now maps plan to its real /v1/plans/health path
(enable id no longer equals route prefix for plan, as is already true for
account/runtime/agent).
cloud#321 landed the Go-only Dockerfile (FROM cloud-flags:latest) but the
release.yml integration wasn't in it, and the reusable could not publish
cloud-flags at all — so every release since has FAILED at
'FROM cloud-flags:latest: not found'. Two fixes:
1. native/flags/Dockerfile base ghcr.io/hanzoai/mirror/rust -> public.ecr.aws
(digest-identical). The hanzoai/ci reusable builds cloud-flags with the repo
GITHUB_TOKEN, which 403s pulling the cross-repo-linked private mirror package;
a public base is GITHUB_TOKEN-pullable, so cloud-flags finally publishes.
2. release.yml resolves console-embed/agent-skills/cloud-flags :latest to
IMMUTABLE digests at release time (crane) and passes them as CONSOLE_IMAGE/
SKILLS_IMAGE/FLAGS_IMAGE build-args to BOTH the smoke build and the push build,
replacing CONSOLE_CACHEBUST. Reproducible (pinned, not floating :latest) AND
fresh (a console/skills/flags change is a new digest). A MISSING artifact FAILS
the release BEFORE build/smoke/tag — the receipt invariant holds, never a
phantom tag on an image that could not embed the real console.
Preserves #322's functional + migration smoke gates (different sections).
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The finance domain had no benchmark; this measures the read that replaced the
commerceinproc self-dispatch — finance.ListUsage over a per-org SQLite ledger,
at 100/1000/5000 seeded debits. Backs the reproducibility claim in the
hanzo-unified-tenant-cloud paper (1.25/10.5/61 ms). Also surfaces a real N+1:
store.Entries fetches postings per row the usage view never uses — a
postings-free read would cut this ~10x (follow-up).
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
Co-authored-by: hanzo-dev <dev@hanzo.ai>
There were two words for one idea. money.Amount.Minor() returns an integer at the
CURRENCY's decimals — money.USD declares 2, so cents. cloudmoney.Amount.Int()
returns an integer at 18 decimals — atto. Both are "the backing integer", neither
name says which, and they differ by 10^16.
That is not a style complaint. It is how
Amount: cloudmoney.FromInt(u.Charge.Minor()), // exact 18-dp USD, no floor
got written, reviewed, and shipped. It type-checked (both sides *big.Int), it read
as "make an Amount from the integer", and it billed a $17.376 zen call as
$0.0000000000000017 until v1.801.44. The package documented the trap in prose —
"Take the decimal, never Amount.Minor()" — because a comment was the only place
the unit existed. Prose is not a type.
So name the unit, not the Go type: FromInt -> FromAtto, Int() -> Atto(),
IntString() -> AttoString(). FromCents/Cents already did this and were never
confused with anything. Now the units are visible at the call site, and the
mistake reads as one: FromAtto(x.Cents()) is obviously wrong where
FromInt(x.Minor()) was obviously fine. It no longer type-checks either —
FromAtto takes *big.Int, Cents() returns int64 — so the pairing that cost us the
money is now two independent kinds of error instead of zero.
Values are untouched: Atto/AttoString return exactly what Int/IntString did, so
the treasury ledger hash and every stored 18-decimal string are byte-identical.
No migration, no data change — only the names, and the compiler found every one
of them (45 sites; a regex could not have, because .Int() also belongs to big.Int
and decimal).
Zero regressions: the failing set is identical to origin/main.
luxfi/keys v1.4.0 was force-moved on the remote (tree 5153d639→80a3745a),
producing a go.sum checksum mismatch that broke `go mod tidy`/`go build` and
the hanzoai/cloud image build in CI. Every keys tag v1.1.0..v1.4.0 was
force-moved with content changes; only v1.4.1 is byte-stable (local==remote
tree 17551acd). Adopt v1.4.1 as the single stable version. Its go.mod floors
the unified luxfi stack, so the transitive set moves forward (geth 1.17.12→
1.20.1, consensus 1.35.32→1.36.2, crypto/database/warp/zap/…), all within v1.
go.mod/go.sum only; diff vs origin/main is luxfi/* exclusively. Embedded
iam2 v0.14.0, apps.go, and concurrent commerce/agent work untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The console SPA, agent-skills catalog, and native flags staticlib are each now
built by their OWN CI as a versioned immutable image and PULLED into the cloud
build, instead of rebuilding node+python+rust from scratch every release. The
console stage (cold npm install + full Next static export, cache-busted every
build) was the ~20-min long pole; it is now a registry pull.
- Dockerfile: console/skills/flagslib stages -> FROM ${CONSOLE_IMAGE}/
${SKILLS_IMAGE}/${FLAGS_IMAGE} prebuilt pulls; COPY sources updated
(/dist, /catalog, /libhanzo_flags.a). Mirror golang+alpine bases, GOPRIVATE,
and every RED gate (SQLCipher proof, modernc guard, cek frozen-format) are
unchanged. Pins are ghcr.io so both buildx lanes pull directly; mirrored to
registry.hanzo.ai (S3) for GET-flow consumers. release.yml still owns the
cloud image + v* tags (it resolves CONSOLE_IMAGE to a fresh console-embed
digest, as CONSOLE_CACHEBUST did).
- native/flags/Dockerfile: the cloud-flags artifact (rust -> scratch /libhanzo_flags.a).
- hanzo.yml: images: cloud-flags (distinct artifact, never a v* tag) + zccache
RUSTC_WRAPPER on the native-flags gate (no-op unless the runner carries it).
Companion artifact publishers: hanzoai/console#(console-embed),
hanzoai/openapi#(agent-skills).
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
Co-authored-by: hanzo-dev <dev@hanzo.ai>
featureflags -> flags (Hanzo's Unleash-analog / hanzoai/flags): ONE runtime
decision engine, (Principal, context) -> verdict, evaluated in-process, hot.
Fold the launch-control waitlist off its DUPLICATE SQLite mode store
(clients/featuregate) onto the one engine:
- a service's waitlist mode IS the platform switch waitlist.<svc>,
evaluated through the same native evaluator as every other flag
- the host->service registry folds into clients/flags (waitlist_store.go,
mode column dropped — the mode is the switch)
- featuregate.Enforce stays but as a CONSUMER of flags.WaitlistModeForHost
(the decide), via an injectable Gate seam
- /v1/featuregate/mode -> served by flags (the decide), same wire shape
- /v1/admin/services -> an admin lens like /v1/admin/flags
- per-user approval still reuses IAM (featuregate/approval.go, unchanged)
- featuregate dropped from apps Wire() — it exposes only Enforce now
The flags package doc names the aspirational end-state: authz (access policy)
and entitlements (product-access policy) are the SAME (Principal,context)->
verdict shape and could COMPOSE this one engine. NOT touched here — flagged only.
Tests (green): featuregate RULE acceptance matrix + approvals; folded-registry
store (Seed/ServiceForHost/List/Upsert); flags env/board/parsers;
apps TestWireOrderMatchesFrozen. Native-engine + cek store tests need
CGO+libsqlcipher+libhanzo_flags (CI), unchanged.
The bridge's mint guard kept its own list of commerce's mint routes under the
comment 'kept in lockstep with api/billing/handlers.go'. A comment cannot hold
two lists together, and it hadn't: 10 paths here against the 16 commerce gates.
Six money-mint routes were outside this guard entirely.
Commerce now DECLARES its gated surface — middleware.Mint records what it gates,
MintRoutes() exports it (commerce v1.49.0) — so the guard reads that declaration
instead of copying it. A mint route added in commerce is covered here with nobody
remembering anything, which is the only version of this that survives contact
with a busy repo.
Registration is what populates the registry, so the guard registers commerce's
billing routes before reading it. The request-level assertion is unchanged and is
the part that matters: an ordinary org user's call must never REACH commerce,
because arriving at all means arriving with the admin service token that
satisfies MayMintMoney.
Proof it bites: allowlisting 'deposit' fails with the escalation itself —
'POST /v1/billing/deposit reached commerce carrying Bearer svc-tok'.
16/16 gated routes refused. All six bridge tests pass.
Every 30s (its own ticker in the connect select loop), scan --studio-dir/output
recursively for image files new or changed since the last scan and POST each to
<studio-url>/v1/library/upload with the worker's bearer, tagged ?node=<identity>
and ?subpath=<subfolder>. So EVERY render lands in studio.hanzo.ai — including
ones produced OUTSIDE the job path: a graph hand-run on the node, or a render
that finished after its activity was reaped (the stranded-late-render class). The
mirror's independence from claims/activities is the point.
State is a tiny in-memory map[relpath]size that skips unchanged files; the studio
endpoint dedupes byte-identical uploads, so a re-scan after restart is cheap and
harmless. One log line per newly stored file; upload failures are summarized once
per scan and retried next tick (no 5xx spam). New --studio-url flag (default
https://studio.hanzo.ai, HANZO_STUDIO_UPLOAD_URL honored) so the mirror works
WITHOUT jobs. No new deps. Test: cli/gpu_mirror_test.go (new/changed uploaded,
unchanged skipped, node+subpath+bearer carried).
Bump github.com/hanzoai/agent v0.1.1 -> v0.1.2 and have the in-process completer
return the agent typed hz.UpstreamError{Status,Body} on a non-2xx completion. The
round then passes a caller-facing 4xx (402 insufficient_balance, 429, 403) through
verbatim so a no-credit user sees "add credits", not an opaque "agent: completion"
502. Proven live: POST /v1/agent with a real hk- key ran the round end-to-end and
the completion returned 402. Dropped the now-unused clip() helper. Simplified
commerce_errorscope_test to the real invariant (typed 403 never flattened to 5xx)
now that commerce v1.48.10 honors status itself; the scope stays as the boundary.
It still said the MountFunc takes the app as `any` and that in-repo subsystems
recover it via Typed. Neither is true: MountFunc names *zip.App and Typed is
deleted. The app is handed to each mount as itself.
CHANGE 1 — the /v1/sync git provider drives GITEA (the one git store), not the
retired cloud embedded store:
- clients/sync/gitea.go: a Gitea REST + go-git client (GIT_ADMIN_TOKEN, GITEA_URL).
Inbound = fast-forward-only go-git fetch(source) -> push(Gitea), so a diverged ref
is a conflict, never overwritten (split-brain guard, now on Gitea). Outbound = a
Gitea push-mirror (sync_on_commit) so Gitea itself propagates every commit. Fails
closed when GIT_ADMIN_TOKEN is unset.
- git_provider.go Reconcile and sync_api.go patch/delete now compose those Gitea
primitives; the cloud embedded seams (InboundGitSync/ImportGitRepo/EnsureGitMirror)
are retired from the sync path. resolve() decision, loop guard, cursor idempotency,
and hop limit are unchanged.
CHANGE 3 — webhook reject parity + a pre-existing root test panic:
- Wrap /v1/git/webhook and the slack(events,commands)/discord/teams/telegram inbound
webhooks in cloud.Terminal so a bad-signature 401 / malformed 400 survives the
commerce /v1 500-flatten (uniform reject codes, matching /v1/sync and
/v1/connector/github/webhook).
- config.go LoadConfig: guard the process-global flag registration with a sync.Once,
so a re-entrant LoadConfig (many test callers in one binary) no longer panics
"flag redefined: enable".
CHANGE 2 (retire the embedded git server) is NOT done: it is still a live dependency.
cloneURL resolves to api.hanzo.ai/v1/git (cloud's OWN embedded server) and the
coding-agent orchestrator clones from uploadPack, pushes to receivePack, and reads the
store via VerifyRef. Deferred — migrate coding to Gitea first.
Tests (CGO_ENABLED=0): clients/sync + clients/integrations green; new
clients/sync/gitea_test.go proves reconcile acts on Gitea and the fast-forward guard.
commercemid.ErrorHandlerJSON() was installed as a /v1 GROUP middleware, but fiber
matches group middleware by PREFIX, not by the handle a route registered on. So on
the shared /v1 it wrapped EVERY subsystem mounted after commerce (projects, agents,
wallets, functions, integrations, marketplace, team, s3, analytics, knowledge,
automations, deploy, billing) and flattened their typed zip.HTTPError (403 "X-Org-Id
required", 400, …) into a blanket 500 — the store envelope always renders 500. The
authenticated release smoke (#322) correctly fails on 5xx, so this blocked every
release since it landed.
Fix stays cloud-side (no commerce dep bump, no money-path churn): commerceErrorScope
guards the envelope by commercePrefixes, so it stays on commerce and every other
subsystem renders its own status via zip default. Pre-commerce subsystems (kms,
o11y) already did; this makes the post-commerce ones match. Verified with the REAL
commerce middleware: /v1/projects 403 (was 500), /v1/store/current keeps the 500
envelope. Regression test added.
MountFunc took `app any` and cloud.Typed asserted it back to *zip.App on every
mount — a runtime check doing the type system's job, with a failure branch that
could not fire because the only value ever passed is a *zip.App. Every subsystem
paid for it: 85 call sites read cloud.Typed(x.Mount) instead of x.Mount.
The reason given was circular. cloud said `any` was load-bearing because an
external module (licensing) exposed func(any, Deps) error; licensing said it used
`any` to avoid an import cycle in pkg/cloud. Each pointed at the other, and the
cycle cannot exist: this package already imports zip (build.go), and zip does not
import cloud. The `any` was justified by nothing.
So name the type. MountFunc is func(*zip.App, Deps) error — what every subsystem
already exported and what licensing's own doc claimed all along. Typed is deleted,
the 85 wrappers are gone, and the three in-repo mounts that hand-rolled the same
assertion (mountZen, mountMetrics, MountO11y) just take the app.
The compiler immediately found what the `any` had been hiding: four mounts still
shaped func(any, ...), one of them across a module boundary. That is the point —
a signature drift is now a build failure instead of a runtime error nobody would
see until a subsystem mounted.
Also here, because the same rip surfaced them:
- iam: v1.31.27-0.20260716191958-4400762928a2 -> v1.31.28, and the replace
pinning a second pseudo-version is dropped. The required pseudo-version named
a commit that no longer exists (it was rebased away), so `go mod tidy` could
not resolve it; v1.31.28 is a real tag and a strict superset of what the
replace pointed at. A version, not a coordinate, and no replace.
- licensing -> v0.1.5, which is where the typed Mount ships.
- cloud.OrgConfig is aliased next to LicenseEntitlement. Both are named by
CommerceClient's methods, but only one was exported, so the exported
interface could not be implemented from outside without reaching into
cloud/types — an omission, not a boundary.
- build_registration_test.go tested Typed and nothing else. "Recovers the
*zip.App" proved an adapter passed through its argument; "fails closed on a
wrong type" cannot be compiled now. What is left is the assertion that a
subsystem signature IS a MountFunc — which the build checks.
No regressions: the failing set is byte-identical to origin/main (11 pre-existing
TestAudit_*).
The CLOUD_IAM_IMPL=iam2 identity fold (clients/iam2, identitySpec) was pinned
to iam2 v0.1.1 — an early cut missing the whole console-parity surface. Bump
to v0.14.0 so the embed actually serves what hanzo.id needs:
- RFC/IETF surface (HIP-0111): OAuth2 code+PKCE/refresh/client_credentials/
password, RFC 8693 token-exchange, 7662 introspection, 7009 revocation,
8414 AS-metadata, OIDC UserInfo, SCIM 2.0 Users
- Casdoor verb-alias compat (transitional cutover bridge) for every verb the
live console/gateway hard-code: users/orgs/apps/providers/roles/projects
- operator bootstrap upsert (IAM CR reconciliation), TOTP MFA enrollment,
organization-scoped projects (ScopeSwitcher)
No wiring change — main's identitySpec + co-mingle iam2server.Mount(app, db)
compile unchanged against the v0.14.0 API. zip already at v1.8.3.
Assisted-by: Claude Opus 4.8 <noreply@anthropic.com>
CHANGE A — first-party vs external route naming. Rename /v1/github-webhook ->
/v1/connector/github/webhook, opening the external-platform namespace
/v1/connector/<provider>/webhook (github now; gitlab/others are sibling literal
routes later, each with its own signature scheme). /v1/git/webhook (first-party
Hanzo Git) and /v1/sync (the bridge) are unchanged. No live consumer breaks:
the GitHub App isn't created yet.
CHANGE B — 500 -> real 4xx on the reject paths. mountCommerce registers
commerce's ErrorHandlerJSON on app.Group("/v1"); that filter rewrites ANY error
a downstream /v1 handler PROPAGATES into a hardcoded HTTP 500. Every /v1
subsystem mounted after commerce is wrapped the same way -- git, sync AND
integrations alike. git does NOT escape it: its isolated tests read 401 only
because they don't co-mount commerce (its bad-sig path is never exercised in
prod, so the flatten went unseen). Add cloud.Terminal, which writes a returned
*zip.HTTPError in-band and returns nil, so the filter's c.Next() sees nil and
has nothing to flatten. Wrap /v1/sync (all verbs) and the connector webhook.
sync no-principal is now 401 (was 403) -- an authentication failure, matching
the webhook's bad-sig 401. The commerce filter is left untouched.
Tests reproduce the /v1 flatten filter and assert: sync unauth -> 401, connector
bad-sig -> 401, malformed body -> 400; connector resolves at the new path and
the old /v1/github-webhook 404s.
The global BillingGate priced /v1/agent/* at a flat 1c via a dead legacy
bot-reverse-proxy rule in DefaultPrice, so every agent request was gated: the
unauthenticated preset/conversation reads hit the balance check, failed closed,
and 503d before the orchestrator handler ever ran. The route was mounted and
winning precedence the whole time — the gate short-circuited ahead of it.
/v1/agent is self-metered: the reads are free and POST /v1/agent bills through
the /v1/chat/completions it runs in-process (gated + metered downstream), so the
edge must price it 0 or double-bill. Drop the legacy branch (and its now-unused
cloudEdgePriceCents const); price /v1/agent and /v1/agent/* at 0. Tests updated.
The spec IS the router, not a description of it: apps.Wire() -> MountAll ->
app.Fiber().GetRoutes(). 983 operations / 692 paths / 109 products, served beside
/zap so ZAP and OpenAPI are two projections of one route table rather than two
sources that can disagree.
A drift test proves the bijection and was proved to fire; it already caught
/v1/pricing-policy and /v1/pricing/policy collapsing onto one operationId.
# Conflicts:
# serve.go
A live authenticated smoke surfaced two production bugs; this fixes both and
adds the durable smoke that now guards every release.
BUG — /v1/billing/usage 500 for a valid caller. usage() proxied
"/v1/billing/usage" through commerceinproc, which re-dispatches BY PATH.
Commerce's own billing routes are behind //go:build cloud and never compiled
here, so the ONLY registration of that path is usage() itself — the S2S hop
re-entered the handler, which self-answered "sign in to view billing". This is
the SAME defect balance() was already fixed for. usage() now reads the usage
ledger DIRECTLY from cloud's own finance ledger (finance.ListUsage → the
wallet→revenue debits RecordUsage wrote), off the self-dispatching hop;
split-deploy falls back to the commerce S2S read, unchanged.
BUG — the balance gate 402'd read-only GETs (fixed in hanzoai/ai, pinned via
the go.mod bump). A $0-balance org could not VIEW its own resources. Fixed in
the ai module's BalanceGateFilter — reads never spend, so GET/HEAD/OPTIONS are
exempt; only writes/metered POSTs gate on balance.
SMOKE — cmd/smoke: a durable, authenticated per-subsystem prober. One
side-effect-free read per core subsystem; a read that 402s (balance gate) or
5xx-es (crash) fails the release. Baked into the image (Dockerfile) and wired
into release.yml as the functional gate after the boot check, so a release can
never ship with chat/billing/projects/kms/... down. The smoke token is
KMS/secret-sourced (never hardcoded); absent → the anonymous matrix still gates
public/authed and catches every 402-on-read / 5xx.
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Store every logged-in identity in ~/.hanzo/identities.json keyed by a stable
<owner>/<name> key (admin/z, hanzo/z), with an Active pointer; mirror the active
identity into credentials.json so every legacy single-file reader is unchanged.
A second login as a different owner for the same email (privilege separation:
hanzo-admin-guard vs hanzo-console) is stored beside the first, not over it.
New: auth list (whoami --all), auth switch <owner|owner/name> (alias use),
logout [<owner>]. Token refresh writes through the store (SaveActive) so the
active identity stays fresh after rotation. All files stay mode 0600.
One credential store, active-pointer switch, backward compatible.
apps.Wire() had TWO {Name:"o11y"} entries: the in-repo read plane
(o11y.MountO11y) and the external hanzoai/o11y module wildcard
(cloud.Typed(o11ymod.Mount)). Fold the module wildcard in as the TERMINAL
sub-mount inside o11y.MountO11y (registered after every specific /v1/o11y/*
route, so Fiber in-order match still gives them precedence), delete the 2nd
Wire entry and the now-unused o11ymod import -> ONE o11y spec.
/v1/o11y/health is preserved exactly: the merged entry keeps OwnsHealth=false,
so the generic always-ok route is registered before MountAll (ahead of the
wildcard) — byte-identical to when the module co-entry, also OwnsHealth=false,
triggered it. Frozen wire order updated (two co-owned rows -> one); the
no-duplicate test no longer needs an o11y exemption.
Zero importers, no cmd/session, absent from apps.Wire(); its /v1/code/sessions
surface was never mounted (dark). Removes session.go, store.go, session_test.go.
The one-binary console (console.hanzo.ai — cloud serves console's static export)
reads org members + projects at /v1/iam/*, but cloud 404'd those (IAM isn't
folded in-process yet) and the old /org/iam BFF proxy is pruned from the static
export. So the browser got the SPA shell (HTTP 200 HTML), which the client
surfaced as "Request failed (HTTP 200)" — the Platform page died on "Could not
load".
iam_edge.go adds an org-scoped reverse edge at /v1/iam/* → the standalone IAM,
mounted when IAM is NOT folded in-process (else that subsystem owns the path — no
double-mount). The org is PINNED to the caller's validated, server-minted
X-Org-Id (never a raw client header) — load-bearing, since IAM's own authz is
permissive on the org-keyed routes, so without the pin one tenant could read
another's projects. A super admin may cross; a tenant may not; writes require an
org admin and must carry the caller's own org. Shares the ONE IAM identity
(iamHost/iamCred) with the API-key resolver (DRY).
Verified: 7 gate tests (pin, cross-tenant refuse, super-cross, allow-list, 401,
write-gate, own predicate) + go build clean.
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
Fixes the live recurring tasksd concurrent-map-writes fatal (hanzoai/tasks#18, v1.51.1). Build green; the failing 'Test' check is the pre-existing repo-wide LoadConfig flag-redefine panic (fails on main + all branches), unrelated to this go.mod-only bump.
clients/platform was the only thing still minting Service CRs; every other
declarer is already App. A role-less App dispatches to the operator's service
profile, the same reconcile the Service kind ran, so a tenant workload carries
over verbatim.
Existing tenant CRs still resolve (App first, Service second). A redeploy patches
the kind it IS rather than minting a twin — two CRs on one name is the
commerce-admin ownerRef flap, not a migration. Teardown deletes BOTH kinds:
either alone rebuilds the app the tenant just deleted, still billing.
113 tests pass, 7 new.
clients/platform was the only thing still minting Service CRs. Every other
declarer is already App: universe git has zero hanzo.ai/v1 kind:Service, the
operator calls App "the sole workload reconciler for the collapsed fleet", and 69
of the 80 live CRs are Apps. A tenant app was the exception for no reason — a
role-less App dispatches to the operator's service profile (controllers/app.rs
`classify("") => Dispatch::Service`), which is the same reconcile the Service kind
ran, so an App carries a tenant workload verbatim.
Nothing mints a new Service CR after this. What already exists still resolves:
reads, patches, scales and deletes walk crGVRs() — App first, Service second — so
the 3 live tenant CRs written before this keep working untouched.
A redeploy of a pre-collapse app patches the kind it IS rather than declaring an
App twin. Two CRs claiming one name is not a migration, it is the commerce-admin
flap: both kinds materialize the same children through the same materializer under
one field manager, so the Deployment's ownerRef just flips between them.
Teardown deletes BOTH kinds. Either kind alone re-materializes the Deployment, so
removing only the one that resolves first would rebuild the app the tenant just
deleted — running, and still billing, minutes after a successful delete.
Sequencing (deploy order matters): operator v0.7.7 carries the Claim guard that
makes App the deterministic owner when both kinds claim a name. This is safe
before it — the no-twin rule means a colliding pair is never created here — but
the guard is what makes an existing collision safe to clean up.
REMOVABLE once no Service CR remains in any tenant namespace: drop servicesGVR
from crGVRs() and the delete-both, and this file is App-only.
Tests: 113 pass; 7 new (kind is App, no twin on legacy redeploy, delete removes
both kinds, idempotent delete, resolution order, absence honest). The 10 tests
that asserted a Service CR is written now assert the kind we write.
TestMigrateOverLegacyPlatformApps fails identically on pristine main (sqlcipher
codec, environmental).
The PaaS control plane serves platform.hanzo.ai and read services only: 69 App CRs
run in the scanned namespaces against 7 Service CRs, so the SUPERADMIN drift board
rendered 7 rows for a 69-app fleet — 1 in production — and /v1/paas/health probed
the Service CRD, found it served, and reported ok over a blind board.
Reads now walk App first, Service second, deduped by name. Writes refuse a
git-declared App (Hanzo CD syncs it with selfHeal, so a patch is reverted on the
next sync) and name the file to commit to; a Service CR still patches.
35 tests pass, 0 fail.
Lifts the unified binary off replicas:1 on DigitalOcean (RWO-only block
storage, no RWX) WITHOUT any shared volume. Each org is pinned by rendezvous
hashing (ha.Owner over the static CLOUD_PEERS ring) to exactly one owner pod;
a request whose org this pod does not own is transparently forwarded to the
owner, so every per-org SQLite store (KMS, finance, and every org-keyed store),
audit append, per-org rate ceiling, and prepaid billing debit runs on ONE pod.
THE INVARIANT — no two pods ever write one tenant's SQLite file — is upheld by
two independent guarantees that compose: per-pod RWO PVC (different physical
files per pod) + org→owner routing (all of an org's writes on one pod).
- shardrouter.go: the middleware. Runs immediately after SanitizeIdentity so it
keys on the VALIDATED, server-minted X-Org-Id (never a client header), hashes
the SAME injective SanitizeOrg slug the on-disk path uses (routing key ≡ file
key), and forwards via a fasthttp streaming proxy (SSE/chat pass through — a
streaming chat still carries a per-org billing debit, so it routes too) with
dial-only retry across an owner's roll gap and a 421 loop-guard on divergence.
Static identical membership rules out dual-owner and routing loops.
- config.go: CLOUD_PEERS (id@addr ring) + POD_NAME (self). Sharding auto-enables
only when >1 peer; Validate fails closed if self is not in the ring or if
embedded iam (non-shardable process-local sessions) is co-enabled.
- serve.go: wires the middleware after IdentityMiddleware; shard-aware boot log.
- audit_serve.go: per-shard audit is automatic on the per-pod PVC; stamp the
shard id on the AU-9 checkpoint stream so the tail-truncation monitor tracks
N heads. writerpin.SingleWriter is correct PER SHARD (each pod sole writer of
its orgs); the writer lease is pod-local under per-pod PVC and stays off.
No-op (byte-identical to today) when CLOUD_PEERS names ≤1 pod. N pinned at 3;
rebalance-on-N-change (a tenant-file move) is the documented follow-up.
Tests prove: exactly one owner per org (deterministic, total, evenly
distributed); all pods agree (no dual-writer); owned served locally, unowned
forwarded to the owner (local chain never runs, hop header set, response
streamed); org-less served locally; loop-guard 421; boot-gate fail-closed.
zen prices every SKU as an exact 18-dp value tagged money.USD. money.USD
declares 2 decimals, so Amount.Minor() rescales that value to CENTS, while
cloudmoney.FromInt reads its argument as 18-dp. Composing the two divided
every zen debit by 10^16: a $17.376 charge debited $0.000000000000001738,
and any charge under half a cent folded to a zero that metering.Record drops
before it reaches the ledger — no debit row, call served free.
The spend gate read the same composition. FromInt(est.Minor()).Cents() is
always 0, and AuthorizeVerdict only compares balance against size when
AmountCents > 0, so the size check was dead code and any org with a positive
balance could draw a request of any size. The dust debits never accumulated,
so the cap could not trip either.
Route the seam through one conversion. credit() carries the exact decimal
across and changes only the minor-unit convention — no rescale, no rounding,
no factor — because the decimal is the value and a currency's Decimals is a
rendering convention. cloudmoney.FromDecimal is the typed counterpart of
ParseUSD that makes this expressible without a string round-trip. nano()
takes the typed Amount rather than a bare *big.Int, so a cents integer can no
longer reach the warehouse fold.
The tests asserted Amount.Int() against Charge.Minor() — the same integer on
both sides — so they proved a round-trip and were blind to the unit. They now
assert against known dollar values built through a different constructor: a
$17.376 charge debits $17.376. Reintroducing Minor()-into-FromInt fails all
of them, including a gate test that admits an over-cap request.
The /v1/billing/* bridge attaches the commerce service token, and that token
satisfies commerce's MayMintMoney. The subpath was checked for traversal but
not against a set, so POST /v1/billing/deposit forwarded a mint for any
signed-in user — pinned to their own subject, which aims the mint rather than
stopping it. Commerce closed this on its direct path the day after the bridge
reopened it, and wider: the bridge needs only a validated principal, never an
admin bit.
The forwardable set is now a per-method table, consulted before the token is
attached, 404 on anything else. Per-method because payouts is a GET read and a
POST mint on one path — a method-blind set hands the mint to every reader.
Nothing reachable today: commerce's mint is not compiled into this binary and
commerce.hanzo.svc resolves here. But devnet runs a real mint-capable commerce
and points cloud-api at it under a second name for the same thing; unifying
those names would arm it. The gate should exist before that cleanup does.
/v1/billing/* forwarded any subpath to commerce carrying the admin
COMMERCE_SERVICE_TOKEN, validated only against path traversal. Forwarding IS
authorization there: commerce gates money-mint on
MayMintMoney = IsServiceToken || IsSuperAdmin, so every forwarded path ran with
platform authority. Commerce 403s an org admin who POSTs /v1/billing/deposit
directly; through this bridge the same person was handed the platform's own
credential, and the subject-pinning aimed the credit at their own account.
Pinning is an IDOR control, not an authority control.
Add billingForwardable: a per-method allowlist of the endpoints the console
actually calls, enforced before the token is attached. An unlisted path 404s.
Allowlist, not denylist — a mint route commerce adds tomorrow is unreachable
with no change here. GET and POST are separate sets because `payouts` is a read
AND a mint-gated write; one method-blind set would hand the mint to every reader.
The POST set holds nothing that creates balance from a client-named amount.
Not exploitable in the current topology: commerce is co-resident in cloud, whose
build never compiles mount.go (//go:build cloud), so no api.Route billing bundle
is reachable and the forward loops back to this bridge. It is one COMMERCE_URL
away from live — devnet already runs a standalone commerce.
Tests: an ordinary org user's deposit/credit/refund/credit-grants/husd/allotment
now never reach commerce; the console's 12 real calls still forward; the store
bridge still cannot tunnel into billing.
zen v1.4.0 bills the prompt cache: cache_read was priced in the catalog,
published in /v1/models, and never charged. cloud pinned v1.3.11, so the
fix could not reach api.hanzo.ai, where the traffic lands.
The module graph cannot move: zen's own go.mod is byte-identical between
v1.3.11 and v1.4.0 (same sha256), and cloud holds exactly one zen edge.
NOTE: this is necessary but NOT sufficient. hmoney.Minor() returns cents
while cloudmoney.FromInt expects atto, so apps/zen.go understates every
debit by 1e16 and the spend gate reads AmountCents 0. Correct pricing is
still zeroed downstream. Tracked separately; that fix is what makes cache
billing real.
The route table gets a third projection. /zap replays the /v1 handlers, the
console renders them, and GET /v1/openapi.json now describes them — all read
from the ONE router after MountAll, so none can drift and none holds a second
copy. There is no checked-in spec file and no second registry.
openapi.Live(app) reads app.Fiber().GetRoutes(true) (fiber's own filter drops
Use() middleware); every other function is pure over that []Route. Each
operation is tagged with its product — the first path segment after /v1/ —
so a CLI can build `hanzo <product> <resource> <verb>` with no judgment.
Reading the LIVE router is the only total source: POST /v1/kms/auth/login is
registered as Group("/v1/kms/auth").Post("/login") and no grep can find it,
and the route set is a function of deployment config, so the document varies
per deployment — correctly. Unauthenticated: it grants no capability, every
route it names stays auth-gated, and `hanzo --help` must build its tree
before login.
The drift guard (cmd/cloud/openapi_test.go) is a bijection over the fully
mounted apps.Wire(): 983 operations / 692 paths / 109 products, every live
route present, no operation invented. Shown to fail on a broken translation
(353 routes reported missing) before being restored.
Honest boundaries, asserted rather than papered over:
- No schemas. The router holds func(*zip.Ctx) error; the request type is a
local inside the handler (var req secretPutRequest; json.Unmarshal(...)),
unreachable by reflection. cloud.Handle[S] is generic over the SERVICE,
not the payload. The path to schemas is zip's typed ops, which today
number zero — which is why zip's own generator emits nothing here.
- No responses block. OpenAPI 3.1 makes it optional; fabricating 200/ok on
~900 routes would assert what nothing knows.
- HEAD/CONNECT excluded — forced, not taste. CONNECT has no OpenAPI field;
fiber auto-generates HEAD in startupProcess(), so including it would make
the document depend on lifecycle stage.
- Catch-alls are opaque: POST /v1/billing/deposit is not a route here.
Corrects LLM.md, which the code contradicted: Wire() does exist (apps.go:188),
MountAll does not sort, and byte-identical patterns MERGE rather than panic.
A high handler count is not a collision — app.Post(path, mw1, mw2, mw3, h) is
one registration with four handlers (apps/commerce.go:151), and 34 live routes
are that shape, so the generator never reads the count.
clients/agent adapter injects the ai completion (in-process, billed) + tools.Default()
into github.com/hanzoai/agent; POST /v1/agent live path. Dispatch resolves the caller
via the ONE canonical tools.PrincipalFrom — no reconstructed principal. Builds green
on the v1.801.35 base. DEPLOY-BLOCKED: go.mod uses a local replace for hanzoai/agent
(CI needs a fetchable release + the local module-cache shallow-clone bug resolved).
cache_read was priced in the catalog and published in /v1/models, but never
charged: on an openai upstream every cached input token billed at the full
`in` rate, and on an anthropic one it left the bill entirely. zen v1.4.0
normalizes both dialects into one tally over three disjoint classes
(fresh + cached + cacheWrite == the whole prompt) and derives once per
response. cloud embeds zen via apps/zen.go, so the fix is inert at
api.hanzo.ai until this pin moves.
zen v1.4.0's go.mod is byte-identical to v1.3.11's, so the require edit and
the two go.sum lines are the whole change: no transitive pin moves, and
nothing else in the graph names zen.
The zen margin test never compiled, so CI has been red on main since it landed:
it parsed prices with shopspring/decimal and handed the result to hmoney.New,
which takes hanzoai/decimal. Two identically-named types, one of them wrong.
vet: cannot use d (struct type "github.com/shopspring/decimal".Decimal)
as "github.com/hanzoai/decimal".Decimal value in argument to hmoney.New
Money has exactly one decimal. A second one that merely LOOKS like it is how a
price silently becomes a different number, which is why the compiler is right to
refuse. Parse with hanzoai/decimal (decimal.Parse — the same call zen itself
prices with); this was the only file in cloud importing shopspring.
Read the debit through Amount.Int(), not Amount.Minor(). The test held two money
types at once: zen's hanzoai/money.Amount (Charge/Cost) has Minor(), but
meterUsage returns cloud's clients/money.Amount, whose accessor is Int() — it
wraps Minor() and returns the identical *big.Int, so the assertions are unchanged.
The tests themselves are worth keeping: they prove the debit is the retail Charge
and never the upstream COGS, that a 3x-margin tier collects 3x, and that a
sub-cent call does not floor to zero. They just could not run.
Not run locally: ./apps links a prebuilt Rust staticlib (libhanzo_flags.a) that
is built in CI, not here. vet — the step CI actually failed — passes, and the
package builds.
GOPRIVATE named zap-proto/*, which is public -- all 55 repos, all 6 modules
cloud needs served anonymously from the public proxy. The namespace that is
private went unnamed: github.com/hanzoai/* (467 private repos incl. ai,
account, commerce, orm, xorm, beego, csqlite). It resolved only by falling
through GOPROXY's direct fallback, and passed checksums only because go.sum
already pins everything.
containment.yml then set GOSUMDB=off to compensate. GOSUMDB does not scope to
a namespace: that disabled checksum verification for every module in the
build, public ones included, in the image that handles payments -- breaking
the invariant the Dockerfile three files away states and honors.
Public modules keep proxy and sumdb immutability; private ones go direct and
authenticated. Verified under the CI and Dockerfile env with -mod=readonly and
sumdb on: build 0, go mod verify all verified.
One reconcile loop, one place a sync happens. clients/sync owns the Sync
record (source/target/direction/trigger/cursor), a per-org sqlite `sync`
table, /v1/sync CRUD + /v1/sync/:id/run, and a kind→Provider registry with a
single Reconcile(sync,event) contract. Git is the first provider, composing the
existing git object-plane seams (InboundGitSync inbound, ImportGitRepo reconcile,
EnsureGitMirror outbound) — no second copy of any git op.
Triggers resolve to Syncs and enqueue the engine (cloud.Sync), never sync
directly: the GitHub App webhook now serves the path it actually fires at
(/v1/github-webhook, was a prod 404 at /v1/integrations/github/webhook) and hands
the verified push to the engine; the Gitea push webhook gains a loop guard
(skip pusher == GIT_SYNC_ACTOR). Loops break on the engine cursor (identical
fingerprints are a no-op) + actor guard; chained propagation (a sync's target is
another's source) is bounded by a hop limit.
Seams (root cloud): SyncFunc + RegisterSync/Sync, GitMirrorController +
EnsureGitMirror. Fail-closed on unmounted engine / missing secret. Tests
(CGO_ENABLED=0): engine loop-guard/idempotency/chain/hop-limit, git resolve,
CRUD+patch+run, webhook signature/isolation/enqueue, gitea loop guard.
GOPRIVATE listed github.com/zap-proto/*. Every zap-proto repo is public — all 55
of them — and every zap-proto module this build needs is served by the public
proxy anonymously. It was never the reason anything resolved direct.
The namespace that IS private went unnamed: github.com/hanzoai/* — ai, account,
commerce, orm, xorm, beego, csqlite and ~30 more. Those only ever built by
falling through GOPROXY's `direct` fallback after the proxy 404'd them, and only
kept passing the checksum step because go.sum already pins them, so no sumdb
lookup happens. It worked by accident, one added dependency away from failing.
containment.yml compensated for that unnamed namespace with GOSUMDB=off, which
does not scope to hanzoai — it disables checksum verification for EVERY module in
the build, the public majority included. The Dockerfile's comment reasoned the
same way inverted ("hanzoai/* and luxfi/* are PUBLIC ... only zap-proto/* is
exempt"); hanzoai/* is largely private and luxfi/* (all 37 deps here) is public.
Naming github.com/hanzoai/* is what the off switch was standing in for. GOPRIVATE
implies GONOPROXY+GONOSUMDB for exactly that namespace, so the private modules go
direct+authenticated and skip the sumdb that cannot see them, while zap-proto and
luxfi keep the public proxy + checksum db — the immutable hashes that make a
force-moved tag unable to break or poison the build. GONOSUMDB and GOSUMDB=off
are dropped: scoped by GOPRIVATE, they are redundant, and blanket-off is a
supply-chain regression in a money image.
hanzoiam/* is not listed: a74b7de reverted the scim/saml/ldap embed, so nothing
in go.mod requires it. It goes back when the modules resolve, named as private.
Verified under the exact CI/Dockerfile env (GOPRIVATE=github.com/hanzoai/* only,
GOSUMDB=sum.golang.org, GOFLAGS=-mod=readonly): build exit 0, vet exit 0,
`go mod verify` = all modules verified, and `go mod download` resolves both a
public module (zap-proto/zip) and a private one (hanzoai/ai).
Runtime routing policy now sources from the GlobalDefaultOwner OrgSettings
row (admin.hanzo.ai-editable SQLite), env demoted to deprecated fallback.
Per-org override unchanged.
The PaaS control plane serves platform.hanzo.ai and read `services` only. The
fleet collapsed onto `kind: App` and this reader never followed: 69 App CRs run
across the scanned namespaces against 7 Service CRs, so the SUPERADMIN drift
board rendered 7 rows for a 69-app fleet — 1 row in `hanzo`, production, and that
row is a duplicate CR. It was not an error anyone could see. The board looked
plausible and was blind to 68 of 69 services; every read of an App-declared
service 404'd; and /v1/paas/health probed the Service CRD, found it served, and
reported ok — status theater over a blind board. clients/deploy already solved
this with an App-first/Service-second read order; this gives the same order to
the reader that needed it.
Both kinds are listed and deduped by name: one workload is one row even when an
App CR and a Service CR both claim the name, because a Deployment has one
controller ownerRef and the operator's Claim guard gives it to the App.
The write path is the harder half. Hanzo CD syncs 68 of the 69 App CRs from
universe `infra/k8s/operator/crs/` with selfHeal on, so patching an App CR here
is reverted on the next sync — the deploy would report success and silently roll
back. That is worse than refusing, so deploy and release now resolve which kind
holds the workload and refuse a git-declared one, naming the file to commit to.
A Service CR has no git declarer (cloud and kubectl write them directly), so it
still patches exactly as before — the tenant plane and the untransitioned CRs are
untouched. release.go's "no ArgoCD" claim was true when written and is not now.
This does not restore an admin deploy button for the git-declared fleet. Under
GitOps that button belongs in git, and ImageUpdate already names that seam
(registry→git→cluster). Whether /v1/paas/deploy should commit to universe or be
retired is a CTO call, not one to make silently inside a read fix.
Tests: 35 pass, 0 fail, including the pre-existing release suite. The fleet-sees-
App-CRs case is seeded to the shape of the real `hanzo` namespace.
Reverts e4b3c88f. go.mod pinned github.com/hanzoiam/{ldap,saml,scim}, but none
of the three repos exist: https, ssh, and API all return not found, and ldap
was never fetched by any proxy or cache. Cold machines — including release
runners — fail go mod download, so main could neither build nor release.
Local builds passed only on module caches warmed 21:36-21:48 UTC today.
iam2 restored to v0.1.1, the state prod v1.801.38 ships. Re-land the embed
unchanged once the hanzoiam repos exist and resolve from a cold GOMODCACHE.
Two functions were called Payer and returned different things: account.Payer
returns the ACCOUNT that pays, principal.Payer returned the ORG whose ledger
holds it. Those are different values on the same request — a person in the shared
signup org pays from account "hanzo/alice" held in ledger "hanzo" — and one name
for both is how the gate came to key the pool while the debit spent the person.
Rename it to what it returns: principal.HomeOrg. An org names a ledger; an
account names a wallet within it.
Fix the last copy of the rule with it. clients/metering is cloud's vendored
metering client, and its IdentityFromGatewayHeaders still hardcoded `user := org`
under the same false premise, while its doc claimed cloud "mirrors" the module
"exactly so every product keys the SAME ledger entry" — a promise two independent
copies cannot keep. Both now call hanzoai/account.Payer, so they agree by
construction and the comment is true for a reason.
Its test asserted the divergence in words — "want hanzo (per-org billing key, not
org/sub)" — which is how a premise outlives the code that disproved it. It now
asserts equality with the rule.
public.ecr.aws (ECR Public) rate-limits anonymous pulls with HTTP 429 on shared
CI runners; a 429 on any base pull aborts the release (a release died pulling
python:3.12-alpine). Repoint all five FROMs to 1:1 linux/amd64 mirrors in our
own GHCR namespace, pinned by digest for immutability:
node ghcr.io/hanzoai/mirror/node:24-alpine@sha256:0cb0e7c3195bce740b6c8d8b27432c92360e3b7f1528087f2c50640b177950c6
python ghcr.io/hanzoai/mirror/python:3.12-alpine@sha256:aa679aa4eed6eb56c1dc6ad3f1b98b7d2d788fd961596779d188fdedad97fb38
rust ghcr.io/hanzoai/mirror/rust:1-alpine3.22@sha256:b348cb409ac0a73de15065997a360063cf87465574a15e3e4469862cb8996f02
golang ghcr.io/hanzoai/mirror/golang:1.26-alpine3.22@sha256:47d47cb5cc3c7dac409dcb6c3a98a6263571218046cd02d709527feef804a77c
alpine ghcr.io/hanzoai/mirror/alpine:3.22@sha256:7c8cb692ae09657cbc4a3f3cbd0e8d5a2690ba38386aaaf252dbb060bf5eb2e6
The four the task named plus alpine:3.22 (the final-stage base — same registry,
same 429 exposure) so no FROM still hits public.ecr.aws. release.yml already
logs the build into ghcr.io (GH_PAT, docker/login-action) before building, so
buildx resolves these private mirrors today with no new plumbing. Only the five
FROM lines + one rationale comment change; nothing else in the Dockerfile.
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
Ships the full router flywheel into the cloud binary: the self-probe
(continuous tagged auto traffic → reward ledger), the fit→gate→auto-deploy
→publish trainer, the routing-latency guard (~0.68us heuristic), and
casibase-chat usage/o11y metering. Also carries the zen warehouse+span
wiring already on main.
Migrates the billing-subject callers (balance.go, account/billing.go +
tests) from the ai/object.Payer that v1.816.0 REMOVED to the extracted
github.com/hanzoai/account.Payer — the same rule in its new home (the
concurrent decomplect). Same replace directive for the force-pushed iam
pseudo-version account v0.2.0 pins.
Claude-Session: https://claude.ai/code/session_018PmFAHZvbBSTsuWyebwMra
The iam2 embed pulls github.com/hanzoiam/{scim,saml} (+ ldap under -tags iam2_ldap).
hanzoiam is a distinct org from hanzoai, so hanzoai/* did NOT cover it — the test
phase would route these private modules through the public proxy/sumdb and 404.
Git auth is already handled by the reusable CI's GH_PAT insteadOf (covers any
github.com private repo the PAT reads).
Last code step of the iam2 migration. clients/iam2 blank-imports the Apache-2.0
enterprise features so their init() self-registers into iam2's feature registry;
iam2server.Mount -> feature.MountAll auto-mounts them under CLOUD_IAM_IMPL=iam2.
ONE mechanism (database/sql driver pattern) — no feature.Register call in cloud,
which (feature.Register appends with no dedup) would double-mount and collide routes.
_ github.com/hanzoiam/scim SCIM 2.0 provisioning (/scim/*)
_ github.com/hanzoiam/saml SAML IdP + SP SSO (/v1/iam/saml/*, /v1/iam/acs, /v1/iam/get-saml-login)
LDAP is GPL-isolated: hanzoiam/ldap links goldap (GPL-2.0), so it is opt-in only via
clients/iam2/ldap_enabled.go (//go:build iam2_ldap). The DEFAULT cloud binary stays
copyleft-free — proven: `go list -deps ./clients/iam2/` has no goldap; only under
-tags iam2_ldap does it pull github.com/lor00x/goldap + hanzoai/ldapserver.
deps: hanzoai/iam2 v0.1.1->v0.1.4; + hanzoiam/{scim,saml,ldap}; go mod tidy.
Without this, deploying the per-org store over live data boots an EMPTY store and
orphans every secret — cloud KMS is the authoritative source the kms-operator syncs
out to every service, so that is a cluster-wide outage. New() now runs a one-time,
writer-only, keyed, FATAL-on-error migration BEFORE serving: it opens the legacy
{DataDir}/kms ZapDB (encrypted at rest with the master key), streams every SEALED
secret (kms/secrets/ prefix; the JSON value carries the full coordinate + AES-GCM
ciphertext + ML-KEM wrapped DEK) verbatim into its per-org SQLite file via store.put
— NEVER unsealing, so no plaintext is exposed and the AAD path-binding survives —
then archives the legacy dir to {DataDir}/kms.migrated so the OS lock is released for
good. Idempotent (prior .migrated marker or absent store = no-op; put upserts).
Tests: migrate roundtrip (opens to same plaintext), no-legacy-store no-op, and the
cross-org relocation defense still holds after migration. Full kms suite green (2
pre-existing admin-edge fails only).
The embedded luxfi/kms ZapDB (a Badger fork) held an exclusive OS lock on ONE dir
for ALL orgs — the hard reason cloud ran replicas=1 behind a cross-process writer
lease. This replaces it with per-org encrypted SQLite ({DataDir}/orgs/{org}/kms.db
via cloud.OrgDB->cek, per-db DEK), which has no exclusive-opener lock, so distinct
tenants never contend and different pods can serve different tenants.
Crypto stays in the client (Seal/Open AES-256-GCM envelope, AAD-bound to the FULL
/orgs/{org} path): plaintext never reaches the file, and a record physically moved
into another org's file still fails to Open (cross-org swap defense preserved).
Adds cloud.PlatformDB for the reserved non-tenant (_platform) partition.
Verified: 54 PASS / 2 FAIL in clients/kms; the 2 fails (admin-edge dualmount/authz)
are PRE-EXISTING — proven failing identically on origin/main. concurrent_open_probe
proves per-org SQLite has no exclusive lock. Build + finance (pure-Go) green.
Branch only — NOT for main/deploy until red review passes.
The gate keyed `user := home` — the org pool, always — on the premise that
prepaid billing is per-org. That premise is false. A person in the shared signup
org holds their OWN account: its members are strangers, not a team, and a shared
org is not a shared wallet. That is what IAM's signed billing_account claim states
and what ai's meter debits.
So the gate authorized against a balance nobody drained. Fund the pool and a
signup-org person still 402s, because their usage comes out of their own account;
fund the person and an empty pool blocks them anyway. Two layers, two answers, one
request.
Resolve through hanzoai/account.Payer — the same function ai debits with, on the
same credential — so the gate and the debit cannot name different accounts. The
premise is removed rather than restated.
The masquerade split is preserved by construction, not by care: the account is
resolved WITHIN the home org, so Account.Org IS the home org and a SuperAdmin
acting in another org still bills their own ledger. A claim naming a foreign
ledger is refused, so it cannot redirect a debit into the org being acted on.
Tests assert against the rule rather than a constant, so they cannot drift the way
the premise did: the signup-org person keys their own account, a real org pools,
the claim wins for a person and a project, and every case is checked equal to what
the debit computes.
clients/iam2 mounts the clean-room zip+orm IAM at /v1/iam when CLOUD_IAM_IMPL=iam2,
else beego, byte-for-byte unchanged. Flag OFF by default → inert. iam2 v0.1.1 seam.
apps.Wire's identity slot now calls identitySpec(): CLOUD_IAM_IMPL=iam2 mounts the clean-room iam2 (zip+orm), anything else — including unset, the production default — keeps the legacy beego Casdoor embed byte-for-byte. The two impls own the SAME absolute prefixes (/v1/iam/*, /login/oauth/*) and cannot co-mount, so exactly one occupies the slot per boot and mount order is preserved. Off by default => completely inert until a canary flips the flag; selection (this) stays orthogonal to activation (cfg.Enabled).
clients/iam2 folds the beego-free Hanzo IAM v2 into the unified cloud binary as an in-process identity plane — the either/or twin of clients/iam. Matches the cloud.Typed contract func(*zip.App, cloud.Deps) error: cloud hands subsystems a cloud.Deps (not an orm.DB), so Mount opens its OWN embedded SQLite ({DataDir}/iam2/iam.db, mirroring the beego embed's {DataDir}/iam layout), seeds config new-only+idempotent from the SAME init_data.json the beego iam uses (non-fatal, honest degrade), then iam2server.Mount registers the whole surface at the canonical absolute paths.
Fail-closed like clients/iam: a store-open or mount failure serves 503 on the identity prefixes while every co-resident subsystem stays up; iam2server.Mount's only panic path (a registered enterprise feature) is recovered in safeMount so it never crashes the shared binary. A TODO marks where the parallel-lane hanzoiam/{scim,saml,ldap} feature.Register lines land.
Pins github.com/hanzoai/iam2 v0.1.1 (seam held stable across the parallel internals refactor); transitive MVS bumps are all patch-level within v1.x (zip 1.8.3, orm promoted to direct, luxfi/crypto 1.20.1, argon2id 1.0.0, pgx 5.9.2). Inert until wired — see the apps.Wire gating follow-up.
The compute-next-version step had the same full-registry --paginate as the tag
step (fixed in f0abd21). It ran first, so it could hang before the build. Bound
it to one page too. Both version scans are now O(1 page), not O(registry).
zen's commerce Meter now also calls ai's TraceServedUsage (recordTrace
WITHOUT recordUsage — the commerce debit stays the ONE billing source,
never doubled), carrying zen's exact per-tier retail (Charge) and upstream
COGS (Cost) folded atto→nano, so zen* traffic in the unified binary lands
in hanzo.cloud_usage + the o11y span plane with TRUE margin instead of
being warehouse-blind. Rides the ai v1.813.1→v1.814.0 bump; balance.go
(+ its drift-guard test) migrated to the renamed Payer/PayerOf API —
same subjects, one rule.
Claude-Session: https://claude.ai/code/session_018PmFAHZvbBSTsuWyebwMra
The embedded zen mount serves the zen catalog in-process and never reaches ai's
pipeToFamily, so zen* calls produced ZERO routing events — starving stats, world,
spark retrain, and /v1/feedback joins. Wire cloud's zen Meter to ALSO write the
RoutingEvent through the ONE shared writer object.RecordFamilyRouting (source="family",
served arm = zen.Usage.Upstream, join key = zen.Usage.ResponseID — the client-visible
response id, new in zen v1.3.11 — tokens + retail cost), fire-and-forget beside the
existing debit. Bumps zen v1.3.7 → v1.3.11 (Usage.ResponseID); ai already v1.813.6
carries object.RecordFamilyRouting.
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
The external Hanzo Git server (Gitea fork, git.hanzo.ai) POSTs push events
here so a push landing on it drives the SAME push-to-deploy core the embedded
smart-HTTP receive-pack path drives: fireBranchBuild -> cloud.OnGitPush deploy
trigger + EmitLifecycle. One code path, no duplication.
HMAC auth: X-Gitea-Signature is hex HMAC-SHA256 of the raw body, verified
constant-time against GIT_WEBHOOK_SECRET (KMS-synced hanzo/prod:/git/webhook-secret).
Fail-closed 401 on unset secret or mismatch. Only X-Gitea-Event: push acts;
others 204. Zero-SHA / non-branch refs are no-ops.
main pinned ai v1.813.6 (which deletes billing_subject.go for the Payer/Account
refactor) without migrating the call sites, so main did not compile:
clients/billing/balance.go:59: undefined: aiobject.BillingSubject
This migrates the call sites to the ONE rule: Payer(Credential)->Account
(owner = Person|Org|Project). Ends the live 402 where a self-serve signup's top-up
minted to the shared org pool 'hanzo' while the gate debited 'hanzo/alice' ($0) --
customer paid, locked out, money in a pool their org-mates could spend.
Green: go build -tags 'libsqlite3 sqlite_fts5' ./clients/... rc=0;
go test ./clients/billing ./clients/account rc=0. go.mod/go.sum identical to main.
main pins ai v1.813.6, which contains the Payer/Account refactor (billing_subject.go
and its PERSONAL_BILLING_ORGS/ORG_BILLING_ORGS lying default are deleted). But the
call sites were never migrated, so main does not compile:
clients/billing/balance.go:59: undefined: aiobject.BillingSubject
clients/billing/balance.go:61: undefined: aiobject.BillingSubjectFromUserKey
Migrates balance.go + finance.go + billing.go to the ONE rule —
Payer(Credential{Owner,Name}).Subject() / PayerOf(org,key).Subject() — and DRYs the
duplicate subject resolvers. This is the cloud half of the fix that ends the live
self-serve 402 (top-up minted to the shared org pool 'hanzo' while the gate debited
'hanzo/alice' = $0).
go.mod/go.sum untouched vs main. Green: go build -tags 'libsqlite3 sqlite_fts5'
./clients/... rc=0; go test ./clients/billing ./clients/account rc=0.
The console top-up (clients/account/billing.go) and the finance read
(clients/billing/finance.go) each re-implemented the billing subject as "always
the org" — the reverted lineage's rule. Against ai's gate, which bills a signup
person per-person, that is the split: money minted to subject "hanzo" (the pool)
while the gate debited "hanzo/alice" ($0) → the paid-up member 402'd.
Delete both twins; resolve the subject through the ONE rule, ai/object.Payer,
keyed on the IAM username (X-User-Name) the gate also keys on. Top-up credits and
console reads now land on the SAME account the gate debits — they cannot drift
because there is one function. The killed PERSONAL_BILLING_ORGS / ORG_BILLING_ORGS
env is inert here too (test proves hostile values change nothing).
Requires ai v1.809.5, which must be tagged FROM ai main after decomplect/account-payer
merges — NOT off a branch. The prior one-rule fix was tagged off an unmerged branch
(v1.806.8/.9), main never got it, and every later tag resurrected the allowlists;
that is why this bug is live. go.sum refreshes via `go mod tidy` once the tag exists.
Verified locally via a replace to the ai branch: clients build + twin tests green.
The repository_dispatch deploy hub is gone (universe image-update.yml
removed in universe d07cf945; the dispatches were silently suppressed by
the flagged sender account regardless). Deploys are declared-tag bumps in
universe crs/, synced by Hanzo CD (ArgoCD, ns hanzo-cd) and reconciled by
the operator. [skip ci]
registry.hanzo.ai/{hanzoai,luxfi,zooai}/ join the /v1/runner push allowlist —
Wave 0 of the native CI/CD migration. Until now only release.yml's crane
mirror could reach the fleet registry; the native BuildKit lane was
ghcr-only by policy.
The "atomic free-version assignment" step paginated every container version
(`gh api --paginate .../versions`) to find the max release number. As the
registry accumulated tags this grew unbounded and hung the step for 30+ min,
livelocking every cloud release. Container versions are created newest-first and
version tags are monotonic, so the max is always on the newest page — query one
bounded page (?per_page=100) instead of the full history. Fast + correct.
Values, not places: the finance-backed credit-ledger adapter is qualified by its
namespace (apps.ledger), not a braided ledgercoreCredit compound. Rename the type
ledgercoreCredit → ledger, the file commerce_ledger.go → ledger.go, and scrub
"ledgercore" from prose (the ledger IS the core — "core" adds nothing). No behavior
change; admin/core tests green, changed packages build.
POST /v1/chat runs one LLM tool-calling round that lets a model manage a system
via tools. Composes existing cloud pieces, reinventing nothing:
- LLM routing + per-org reserve/settle billing: the ai subsystem's
/v1/chat/completions, invoked in-process (Fiber Test) — the only path that
both returns tool_calls AND carries the billing gate.
- tool plane (clients/tools): the org's registered MCP/registry tools are
offered to the model and dispatched server-side (activation + price gated).
- capabilities: graph (advisory node-ops -> ops the client applies) and create
(server-executed, tools = the org's registered MCP render services).
Returns {reply, actions, ops}. chat mounts before ai so /v1/chat resolves here
(the ai /v1/chat alias is shadowed); ai keeps /v1/chat/completions.
Carries the PAID-Enso revert + the family learning loop: per-family-call RoutingEvents
+ shadow A/B (records what the learned engine would have picked), /v1/feedback signal
contract (up/down/regenerate/switch/abandon/accept/revert/rating/dismiss) with the
online reward forward to the engine's /route/observe, ROUTER_ADMIN_TOKEN service-auth
on the training-data exports, and Zen/Enso provider branding. Lights up /v1/router/stats
+ world.hanzo.ai (shadow-vs-served agreement).
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
Ships the per-org router policy (get/update-router-policy + the org > '*' >
conf fold on every auto route), the routing-reward ledger, the nano margin
ledger (costNano/billedNano/marginNano + unpriced flagging), and the
gaps-only DigitalOcean usage backfill (POST /v1/admin/usage/backfill-do,
dry-run default). The release image also re-embeds console@main
(CONSOLE_REF=main), picking up the console Router page (v8.4.137+).
Claude-Session: https://claude.ai/code/session_018PmFAHZvbBSTsuWyebwMra
The last mile of "one way to grant credit": cloud implements commerce's
creditledger.CreditLedger (v1.48.5) over the native finance ledger and injects it
at mountCommerce (EmbedConfig.Ledger). Now commerce's POST /v1/billing/credit AND
the admin.hanzo.ai grant (/v1/admin/customers/:org/credit) both mint into the SAME
per-org finance wallet the ai prepaid gate reads — a granted credit is immediately
spendable, no split ledger. The admin path drops its parallel finance.Deposit for
the one creditledger.Credit call (idempotency key rides through; commerce HTTP
deposit remains the split-deploy fallback).
- apps/commerce_ledger.go: ledgercoreCredit adapter (compile-time asserted against
commerce's exported interface); org-pool wallet (Subject==Org); fails closed with
no co-resident finance.
- apps/commerce.go: inject Ledger: ledgercoreCredit{} at mountCommerce.
- clients/admin/core/grant.go: grantDeposit prefers creditledger.Get() (the one
ledger) over its own finance.Deposit.
Verified: changed pkgs build green on commerce v1.48.5 + ai v1.812.0; admin/core
grant tests pass. Ships with the ai free-tier removal (v1.811.0+) already on main.
One secret id everywhere: BuildKit's gitsource convention (GIT_AUTH_TOKEN,
which the fabric's buildFrontendCmd already attaches) is also what the
Dockerfile's console-embed and go-mod fetch stages mount. gh_token was a
second name for the same credential.
The buildkit Job now surfaces the console-git-token Secret (optional) as the
GIT_AUTH_TOKEN env, and buildctl attaches it as the same-named build secret;
BuildKit's gitsource presents it as the HTTPS credential for the git context.
Fixes the fabric's inability to build PRIVATE repos (github.com/hanzoai/cloud
itself: 'could not read Username' — the rel_HvInGKU failure). Public repos and
Secret-less clusters fetch anonymously exactly as before.
A repo gains a visibility bit: PATCH /v1/git/repos/:name {"public":true}
(or create-time "public"). Public grants ANONYMOUS upload-pack/info-refs
only — receive-pack and the whole control plane stay org-authed, and a
private or missing repo answers the same uniform 404 so anonymous probing
cannot enumerate.
This is what lets the credential-less buildkit build Job (launchDirectBuild
git context) fetch from the embedded git server — the same reason public
GitHub repos build with no env. Private-repo builds remain a later
GIT_AUTH_TOKEN feature.
resolvePackRepo grows the allowPublic branch (fetch-side only): anonymous
callers address org-level repos by the orgRE-validated :org path segment;
the path-vs-identity guard for authenticated callers is unchanged.
Wire a REAL smart-HTTP git push through the actual production seam
(cloud.RegisterPushBuilder <-> cloud.OnGitPush) into the real platform
push builder, and assert the matching app's build is enqueued (a
building deployment for the pushed commit). The two halves were covered
in isolation (clients/git TestPushFiresBuildTrigger; platform
TestBuildFromPush_LaunchesMatchingApp) but nothing connected a live push
to the real builder end to end. Pure-Go dev build (CGO_ENABLED=0).
Deletes clients/bot — a place-name that had collected three concerns (a mount, a
transport, and two unrelated domain wire protocols). One value per module:
run -> clients/bots /v1/bots (no store)
dispatch -> clients/coding /v1/coding-tasks
transport -> clients/runtime address, identity, framing, cleartext policy
machine -> clients/visor /v1/compute/bots (moved off /v1/bots)
session -> clients/agents the one registry (reverted to main)
Fixes a live route collision: visor's listBots silently won GET /v1/bots (the
router merges byte-identical patterns, first wins), so cloud served machine rows
as runs and orgs' real runs were invisible.
POST /v1/bots/run now returns 501 instead of charging $1.00 for a bot that never
booted -- the runtime has no launch endpoint. Stop fails closed: a bare 404 is
502, only a structured error body means already-stopped.
Net -613 lines. runtime never imports bots/coding; the reverse is a compile-error
cycle, so the one-way dependency is structurally enforced.
Cloud's /v1/billing/balance proxy was the only handler registered (commerce's
GetBalance is behind //go:build cloud, never compiled in) so it called itself,
depth 2. Reads the finance ledger via the existing finance.Current() seam and
single-sources the subject from aiobject.BillingSubject so console and gate
cannot drift.
Red found the substrate was wrong: cloud kept its own registry of runs. It minted
an id the runtime had never heard of, so /v1/bots listed runs that did not exist,
stop closed records for runs that were never started, and run charged $1.00 for a
bot that never booted — the runtime has no launch operation at all, and nothing
in run ever contacted it. The registry that exists is the runtime's own tenant
store; it is the only thing that knows whether a sandbox is alive.
So cloud owns policy and the runtime owns the run:
- GET /v1/bots and POST /v1/bots/:runId/stop proxy the runtime, gated by the
validated principal and org. The org is cloud's, never the client's, and the
runtime keys every run under tenants/{org}/ — so a foreign run resolves under
the caller's org, where it does not exist, and 404s.
- POST /v1/bots/run returns 501. There is no launch operation to call, so the
honest answer is that it is not implemented. It no longer charges.
- The agents session plane is reverted verbatim: no surface column, no Agent
filter, no SessionOpen, no in-process stop. It never should have carried a
second copy of the runtime's state.
Absence is only meaningful from a callee that could have said otherwise, so the
transport separates ErrNotFound (the operation ANSWERED absent) from ErrNotServed
(no such operation). A runtime without the stop route reports absent for every
run; treating that as "already stopped" made a stop that cannot fail. It is 502.
Decomplect the transport: clients/bot was named for the host it dials and braided
three concerns — an ops face plus two unrelated domain protocols. It is now
clients/runtime, a domain-free transport (address, identity, framing, cleartext
policy, /v1/bot/* ops face) that must not import bots/coding. Each domain owns its
own wire stub (bots/wire.go, coding/task.go), so the ZAP swap per HIP-0106/0120 is
a seam swap. Wire spec bot -> runtime; cmd/bot -> cmd/runtime.
The duplicate-route guard was a no-op: the router MERGES byte-identical patterns
into one route with chained handlers, so counting entries never saw the collision
it guarded. It now asserts one handler per route, and a test proves it fires on
the original bug. Fiber's Test() defaults to a 1s wall-clock deadline, which made
the isolation guards flake under load; they now pass an explicit timeout.
Route table unchanged: /v1/bots (runs), /v1/compute/bots (machines), /v1/bot/*
(runtime ops).
The console repo-browser (hanzoai/console products/git) needs machine-readable
refs/tree/blob/commits/readme, but git only served those as HTML (ui.go) + the
repo-CRUD control plane. Add the JSON twin, reusing ONE set of go-git read
helpers so the HTML + JSON surfaces can never drift:
GET /v1/git/repos/:name/refs → { branches, tags, default }
GET /v1/git/repos/:name/tree?ref&path → { entries:[{name,path,type,size,mode}] }
GET /v1/git/repos/:name/blob?ref&path → { path,size,encoding,content,binary,truncated }
GET /v1/git/repos/:name/commits?ref&path&limit → { commits:[{sha,shortSha,message,author*,date}] }
GET /v1/git/repos/:name/readme?ref → { path, content, encoding }
- browse.go: the handlers + DTOs (mirror the console GitApi normalizers verbatim).
ref+path ride as ?ref=&path= query params (the UI's own convention) so a slashed
branch is unambiguous. Reuses org()/findRepo()/openGit()/resolveRef()/
cleanTreePath(). Org-scoped (X-Org-Id); a repo outside the caller's org 404s.
Distinct trailing segments — never shadow the :org/:repo smart-HTTP routes.
- ui.go: findReadme → readmeAt (returns filename+content); the ONE readme scan now
backs both the HTML repo home and the JSON /readme (DRY, no duplication).
- browse_test.go: seeds a nested tree + README + binary file, asserts every endpoint
+ org isolation + honest 404/403 + empty-repo refs.
CGO_ENABLED=0 go test ./clients/git/ green (full package); go vet + build clean.
Stays v1.x.x (Go module). Ships with console@main on the next cloud release.
GET /v1/billing/balance and /v1/finance/balance proxied to commerce at the SAME
path they are registered on. commerceinproc publishes the shared zip app and
re-dispatches by path, and commerce's own /v1/billing/* routes are never
registered in this binary (api.Route runs only from commerce's mount.go, which
is //go:build cloud; cloud ships -tags "libsqlite3 sqlite_fts5"). So the proxy
re-entered itself, hit its own sign-in gate with no principal, and reported
"billing upstream status 500".
Co-resident, the prepaid wallet is cloud's own finance ledger: wireFinance points
the ai gate's balance read at it, the edge meter debits it (metering.fetchAvailable
already resolves finance.Current() first), and an admin grant credits it
(core.grantDeposit already prefers it). Read it directly through that same seam.
The commerce S2S read stays as the split-deploy fallback.
The subject comes from ai/object.BillingSubject — the function the ai prepaid gate
itself resolves — instead of a re-implemented copy. cloud and ai each keeping their
own copy of that rule is what let them drift: the console scoped to the org while
the gate scoped to "org/user", so the view showed a funded org while the gate
refused the member.
Fail posture unchanged: a balance that cannot be read is UNKNOWN and surfaces as
502 — never rendered as a zero balance. The sign-in gate and org scoping are
untouched; the org still comes from the validated principal only.
Tests: self-dispatch pinned at depth 2 (the seam's real mechanics, which the
existing SetHandler-stub test cannot see); router semantics probed (byte-identical
patterns MERGE and silently shadow; equal-specificity param-name conflicts PANIC at
registration; most-specific wins over registration order); balance regression,
gate-subject parity, unreadable-is-not-zero, sign-in gate, and cross-org isolation.
A wire adapter, not a second pipeline: POST /v1/insights/e accepts PostHog-
shaped payloads (single or {batch}) from @hanzo/insights and any compatible
SDK, maps $-properties onto the native CaptureEvent, and rides the SAME
capture path (tenant gate -> normalize -> scrub -> hanzo.events). GET
/v1/insights/events is the console's tenant-scoped recent-events read; GET
/v1/insights/health reports the surface. Flags deliberately stay at /v1/flags.
Scale path stays stateless: accept on any replica, pooled batch INSERT sink;
the queue-buffered (mq/pubsub -> Datastore consumer) upgrade swaps the exec
behind buildEventsInsert with no handler changes.
The published contract has no pagination, so the list is capped either way; set
it at the store maximum here rather than inherit the store's 100-row default.
GET /v1/bots was registered twice: clients/visor (bot machines) and clients/bots
(bot runs). The router resolves byte-identical patterns by first-registration
without panicking, and visor mounts first, so visor's machine list answered the
console's run list and clients/bots.list was unreachable. The console normalized
machine rows into run rows, yielding one blank-runId row per kind=bot machine
with a dead sessionUrl, and hid the org's real runs.
Name the values apart, one home and one route namespace each:
- bot run -> clients/bots /v1/bots (unchanged; the console + CLI contract)
- bot machine -> clients/visor /v1/compute/bots (moved; a machine is compute)
- bot runtime -> clients/bot /v1/bot/* (passthrough for runtime-owned ops)
Make the run control plane native. clients/bots holds no store: a run is recorded
on the agents session plane under agent label "bot", so the run id is the session
id and one registry serves every kind of agent work. list reads it org-scoped;
stop resolves (org, runId) against that record and drives the runtime only after
ownership is proven, so a run of another tenant is a 404 the runtime never hears
about. An unreachable runtime is a 502 with the record left live.
Two seams (Runs, Runtime) are injected in adapters.go, the only file in
clients/bots importing agents/bot, so handlers unit-test against fakes.
agents: sessions carry a surface (the modality a session runs on, matching the
runtime's own origin.surface) via the established additive-column migration;
SessionFilter gains Agent so a product face reads only its own sessions;
OpenSession takes SessionOpen; StopSession is the single-session twin of
StopSessions and shares its one write path (stopOne).
Contract change: a run id is now the session id, not bot_<hex>. No bot_ id was
durable anywhere -- the old run handler minted an id and stored it nowhere -- so
there is nothing to migrate. Ids stay opaque to clients.
Lands the Enso limited-preview gating into the live api.hanzo.ai binary
(cloud embeds ai in-process via ai.Mount). v1.809.2 adds ai's Enso family
routing (ENSO_URL -> enso pod) + waitlist gating: ModelAccess visibility in
/v1/models, 403 request-access for ungated SKUs, comped bypass of the balance
gate for granted callers, POST /v1/models/:model/access, org `hanzo` seed.
Independent of the concurrent zen v1.3.0 -> v1.3.7 bump (#307), which this
branch is based on top of.
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
The cloud binary serves the zen* family IN-PROCESS via zen.Mount, so the
embedded module version — not the zen pod image — is what serves zen5. v1.3.0
sized the ladder rung by the byte estimate alone: a ~230K-token prompt with a
32K max_tokens budget estimated ~258K (under glm-5.2's 262144 cap), stayed on
glm-5.2, which then saw prompt+output = 262145 and 400'd. v1.3.7 sizes the rung
by need() = messages + tool schema + max_tokens, overflowing a >262144 total to
the 1M deepseek-v4-pro rung. Brings the Enso family + gating in-binary too.
Claude-Session: https://claude.ai/code/session_015Z1iLf7QBrq1LhignJrzDw
The hanzo-flags Rust staticlib compiles with unwinding (panic-guarded FFI);
its _Unwind_* references resolve from libgcc_s, which the alpine runtime did
not carry — /cloud failed relocation at exec ('Error relocating /cloud:
_Unwind_GetIP: symbol not found') and smoke red-gated the release. Add libgcc
to the runtime apk set.
native/flags: hanzo-flags, a stateless Rust staticlib with PostHog-compatible
evaluation — the exact Insights rollout hash (sha1 first-15-hex / LONG_SCALE,
pinned by test vectors), the full vendored property-operator set (exact/regex/
semver/date/relative-date...), condition groups (variant-override groups first),
multivariate cumulative selection, payloads. Pure (defs JSON, ctx JSON) ->
response JSON behind a panic-guarded C ABI: hanzo_flags_evaluate/_free.
65 tests green.
clients/featureflags becomes the NATIVE engine (no external evaluator, no KV,
no network): definitions in per-(org, project) SQLite via cloud.OrgDB
(encrypted at rest via cek), hot in-memory platform snapshot (TTL 15s),
evaluation over cgo. The INSIGHTS_FLAGS_URL HTTP proxy is gone. Bool/Int/
String and the admin Board keep their shapes; sources are flags -> env ->
default (first cockpit write creates the definition — env fallback intact
until then, zero regression). engine_stub (!cgo) degrades loudly, fail-safe.
/v1/flags (org-scoped via principal, project via X-Project-Id): POST /v1/flags
(+/decide alias) evaluate; GET/PUT/DELETE /v1/flags/defs[/:key]; GET
/v1/flags/activity; GET /v1/flags/health. PUT /v1/admin/flags/:key (SuperAdmin)
is the cockpit write path through SetPlatformSwitch — flips apply immediately
in-pod, peers converge within one TTL.
Build: Dockerfile flagslib stage (rust:1-alpine musl staticlib, --locked) copied
to the exact ${SRCDIR}-relative cgo link path; make native; hanzo.yml
native-flags step + clients/featureflags in the hermetic unit gate.
go test ./clients/featureflags ./apps green (FFI live).
The console SPA shipped in the unified cloud binary is a STATIC export of
hanzoai/console whose <title> is baked to the default (Hanzo) brand at build
time. The embed cannot read the request Host, so every host — including
console.lux.cloud and console.zoo.cloud — served "Hanzo Cloud Console" in the
browser tab, a white-label violation. (The standalone Next.js app is host-aware
via generateMetadata, but it is retired from this serving path: cloud serves
the console in-process from the go:embed static export.)
Rewrite the SPA shell's <title> at the serving layer (serveIndex/indexFor) to
the request host's brand, reusing the existing brands registry (BrandForHost):
console.lux.cloud -> "Lux Cloud Console", console.hanzo.ai -> "Hanzo Cloud
Console" (unchanged), console.zoo.cloud -> "Zoo Cloud Console". Matches
hanzoai/console's own `${brandName} Console` output — one source of truth. The
Hanzo/default host returns the embedded bytes unchanged (no regression); a shell
with no <title> is never altered.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Brings live: get-cloud-usages Bearer + balance-exempt (usage panel works at
$0), the /v1/ai/connections usage-import endpoint. Console changes ride the
same release via CONSOLE_REF=main embed. Build green.
Global SQLite registry {service, hosts, waitlistMode} + admin board/toggle
(/v1/admin/services*, /v1/featuregate/mode) + native Enforce middleware and
IAM-backed per-user approval resolver. Ported from recover/featuregate; dropped
the init()/cloud.RegisterWithShutdown self-registration in favor of the explicit
apps.Wire() MountSpec (order after admin, before tasks) + frozen wire_test row.
Distinct from clients/featureflags (a global read-only Insights waitlist_open
switch); this is the per-host launch lever with in-binary enforcement.
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
Two co-residence fixes surfaced by the live founder-journey e2e:
1. /v1/platform/projects (authed) returned 500 'runtime error: invalid memory
address' — iamobj.GetProjects/AddProject deref the package-global ormer.Engine,
nil until the co-resident IAM subsystem initializes it. A nil store must fail
cleanly: the iamStore() guard converts a nil-deref into a typed 503. (The store
is initialized once iam co-mounts — v1.801.21 embedded-SQLite isolation.)
2. Bump commerce v1.48.1 -> v1.48.2: ErrorHandlerJSON preserves a downstream
typed HTTP status, so a co-resident subsystem's 401/403/404 (e.g. platform's
'X-Org-Id required') is no longer flattened to 500 by commerce's /v1 mount.
Regression tests: nil-store -> 503, success + real-error passthrough.
The ValidatorRegistry is the trusted key source for the whole cert, so key binding must be rogue-key-resistant BEFORE the driver sources idVerifier into the cert verifier. Register is now first-writer-wins: a node already bound to one ML-DSA key may not be silently re-bound to a DIFFERENT key (ErrIdentityKeyConflict); idempotent re-register of the identical key stays allowed. The ONLY sanctioned rotation is Rekey(node, newPub, authSig) — self-authorized by a signature under the node's CURRENT key over rekeyTBS under a DISTINCT rekey context (not popContext/certContext); an attacker without the current secret key cannot rotate. This is the registry action a consensus-ordered OpRekeyValidator applies (op plumbing lands with the RSM rewire).
5 R1 tests green; existing OnePodOneShare + full suite unaffected.
Finding 1: guardBFTFloor(len(keys), quorumWeight) fails closed unless the floor is the byzantine-safe BFT quorum (2n/3+1) for the validator-set size — so a mis-wired wallet-custody t (=3 for n=5) can neither compose nor verify a cert even if a future driver passes it. Called at the top of ComposeControlPlaneCert and VerifyControlPlaneCert. This is the exact two-threshold trap inc-1 hit (cluster.go wiring Pulsar threshold = quorum), now impossible in the cert core.
Finding 3: three raw-craft verify-side tests drive hand-built malicious certs (bypassing the honest composer) straight at VerifyControlPlaneCert — sub-quorum threshold-lie (ErrQCThresholdBelowFloor), attacker-key stuffing (ErrQCMerkleInclusion), weight inflation (ErrQCAggregateWeight) — plus a self-guard test (t=3 refused). 14 cert tests green.
Flag still false; ceremony untouched; existing byzantine suite green.
The shipped Gen-3 weighted-quorum cert the design chose over the blocked threshold-Pulsar path. Each pod signs the canonical quorum message INDEPENDENTLY with its seam-a ML-DSA-65 key under a DISTINCT cert context (RED R4: certContext != popContext); the cert is a quasar.ConsensusCert carrying one EvidenceWeightedSigSet leg (a WeightedQuorumCert of N independent FIPS-204 sigs + weighted-Merkle quorum). Verification is quasar.VerifyConsensusCert under a control-plane ConsensusCertPolicy that requires the weighted-sig-set PQ leg — the shipped, audited verifier; NEVER the structural QuasarCert.Verify. No DKG, no threshold aggregate, no unshipped luxfi/pulsar core: soundness rests only on stock FIPS-204 verify + the weighted-validator-set Merkle commitment. Composes against luxfi/consensus v1.35.32 (BuildWeightedValidatorSet / BuildWeightedQuorumCert / VerifyConsensusCert).
VerifyControlPlaneCert binds the cert to the caller's expected position before the cryptographic verify (VerifyConsensusCert pins validator-set+policy but not the caller's height/round/block).
10 standalone crypto tests green: real cert verifies under policy; below-quorum / missing-leg / forged-sig / rogue-signer / wrong-position / wrong-validator-set all REJECTED with exact typed errors; R4 proven (a popContext sig is rejected as a cert sig); deterministic composition. Self-contained (does not touch the ceremony); existing suite stays green. Flag NOT yet flipped — ceremony rewiring + the ProductionBCCSigningReady() flip + stub deletion follow.
The ArgoCD fork now lives at hanzoai/deploy (hanzoai/gitops redirects to it), so
the cloud CD control plane takes the matching single-word name: clients/gitops ->
clients/deploy, /v1/gitops/* -> /v1/deploy/*, and the subsystem name (Wire entry,
enablement key, health surface) is 'deploy'.
Same plane, same ArgoCD-grade fleet view over our own App CRs — applications,
resource tree, live manifest + diff, logs, sync, rollback.
Completes the per-app-binary invariant: connectorruntime (HIP-0126) is in apps.Wire() but its generated cmd stub was dropped by a stale reconciliation merge. go generate ./apps restores it. cicd green.
The subsystems/ -> apps/ rename merge (345264d2) missed hanzo.yml's test
gate, which still ran `go test ./subsystems/` — a directory that no longer
exists — so every main CI run failed "./subsystems [setup failed]". Point
it at ./apps/ (the renamed composition root). Unbreaks the release gate.
Red LOW-2: body.Spec.GPUs was only truncated post-decode by Sanitize
(cap 32). register + PATCH now return 400 for a GPU array larger than the
cap instead of silently truncating, so an absurd payload is refused up
front and the client learns the bound. The 4MiB body limit already caps
the transient decode allocation; a normal-sized list still registers.
The link app (unified AI login-manager registry, /v1/links) landed in
apps.Wire() without its generated cmd stub. Regenerating reconciles the
per-app-binary invariant: every Wire() app builds standalone via
apps.ServeSingle AND mounts into the unified binary. Idempotent.
A linked computer now reports what it IS (os/arch/cpus/memory/gpus, "spec")
and what it is DOING now (loadavg/memory/gpu-util, "metrics") so mission-
control can show which machine an agent runs on and whether it can take more
work — without copying the fact onto every session.
- targetspec.go: Spec/Metrics/GPU value types, JSON column codecs, and a
total Sanitize that bounds strings/counts/sizes and coerces floats finite,
so a hostile client can neither bloat the row nor smuggle a NaN/Inf.
- agent_targets gains spec/metrics/metrics_at (crashloop-safe addColumns,
PRAGMA-guarded, no new index); register + PATCH accept them; a metrics
PATCH is a heartbeat and the server owns the staleness clock (a client
can't forge At).
- register upserts by (org, host) so re-linking the same machine refreshes
one target instead of piling up duplicates.
Tests: sanitize bounds, store round-trip, GetTargetByHost org-scope, HTTP
capability+heartbeat, upsert-by-host. Existing target tests stay green.
The feat/link-registry merge (669fee5) added the `link` mount to Wire() —
{Name:"link", after agents} — but never updated the frozen mount-order fixture,
so TestWireOrderMatchesFrozen failed on main (84 specs vs 83 frozen). This merge
inherited that break; adding the missing frozen entry (link, no health, has
shutdown) at its Wire() position makes apps green again.
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
A link revoke tore down live sessions by matching only {org, host,
provider, account}. Those fields come from a link row the caller sets at
upsert, so any org member could stop another member's sessions by
registering a wildcard link (e.g. provider-only) and revoking it.
Make Actor mandatory in agents.SessionMatch and AND it into the query.
The link adapter derives it from the revoking user's subject
(agents.BillingActor), the single place that binds a stop to the caller's
own sessions, so a stop/count can only ever reach that user's sessions.
A match with no actor stops nothing (fail-closed).
Tests: agents.TestStopSessions_ActorScoped (wildcard, host-forge,
no-actor fail-closed, count scope, own-account teardown) and
link.TestRevokeCannotStopCoTenantSessions (full stack through the real
adapter, co-tenant survives). TestRevokeStopsSessions stays green.
The org+user-scoped registry of which provider accounts (Claude Max, ChatGPT
Plus, a Hanzo/api key) a developer has signed into, on which machines, with each
account's latest usage snapshot — the cross-machine view console renders and the
source the redundancy route policy reads.
- clients/link: the Link atom (no secret — metadata + usage snapshot only),
per-org SQLite store (org+subject leading-bound, upsert-on-identity, revoke),
the /v1/links surface, and a pure RoutePolicy (Plan) that orders a user's linked
accounts for redundancy (two Claude Max, then the metered API backstop) carrying
the billing mode per candidate. The store holds no metering client — a
subscription's usage is metered for visibility only and never charges commerce.
- clients/agents: a session now carries the linked account it ran under
(Provider/Account tag), and StopSessions/CountActiveSessions expose the
in-process action a link revoke takes to stop the sessions under a revoked
account/device. Backward-compatible session migration (addColumns).
- subsystems: mount link after agents so a revoke can stop its sessions.
Org+user fail-closed isolation, subscription-vs-api-key billing distinction,
and revoke-stops-sessions are all tested (build/vet/gofmt clean; race+CGO green).
Pulls the DeepSeek <think></think> strip into api.hanzo.ai: reasoning-inlining
upstreams (zen5-pro/zen5-flash → deepseek-*) no longer leak the </think> template
token into the visible answer via the Anthropic-translation path that `hanzo code`
uses. Also carries the hk-key 402 tenant-gate fix. Builds clean (server + hanzo CLI).
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
Each app now builds as its own standalone binary AND still mounts into the
unified cloud binary — one source of truth (apps.Wire()). Two pieces:
1. apps.ServeSingle(name) — the ONE way to run a single app standalone: validate
the name against Wire(), then cloud.Serve(Wire(), []string{name}). It is the
path cmd/hanzo's 'hanzo <name>' already uses; promoting it to apps (which
already imports cloud, so no cycle) lets every cmd/<app> stub reuse it instead
of re-implementing the dispatch. Adding an app in Wire() is still the one edit.
2. cmd/gen-app-cmds — a tiny generator (go:generate directive in apps/apps.go)
that parses Wire()'s {Name: "..."} literals and writes one cmd/<app>/main.go
stub per app (apps.ServeSingle("<app>")). Generated, not hand-maintained; a
re-run is idempotent (writes only on content change). 80 unique apps today.
Verify: go build ./apps/... green; apps.ServeSingle added; go generate ./apps is
idempotent; go test ./apps/... green (TestWireOrderMatchesFrozen at 84 specs); a
sample of generated cmd stubs (kms/account/agents/platform/storage/audit/
analytics/ads/ingress/billing/o11y) build standalone green. The full ./cmd/...
build is heavy (80 binaries x the cloud tree) and needs CI parallelism; the
stubs themselves are sound.
The cloud binary mounts ~84 'subsystems' into one unified binary; they read as
apps, so the package follows. Purely internal: dir subsystems/ -> apps/, package
'subsystems' -> 'apps', the one file subsystems.go -> apps.go. Three import sites
updated (cmd/cloud, cmd/hanzo). Broken path/filename references in comments
(subsystems.Wire, subsystems/commerce.go, subsystems.go) -> apps.*; the package
doc + LLM.md/docs paths follow.
No external breakage: the package is same-module (never in go.mod), no external
module imports it. The frozen-order test TestWireOrderMatchesFrozen compares
MountSpec.Name strings, not the package name — passes unchanged at 84 specs.
Standalone via 'hanzo <name>' and per-app cmd binaries are unchanged by this
rename (D2 adds the cmd binaries on top).
Verify: go build ./... green; go vet ./apps/... ./cmd/... clean; go test ./apps/...
green; TestWireOrderMatchesFrozen green. cmd/cloud TestMountAllAndServeHealth
fails identically on clean main (needs CLOUD_KMS_MASTER_KEY_REF), not this change.
migrateSessions created ix_sessions_org_target and ix_sessions_org_host in
the table DDL, which runs before addColumns adds target/host to a
pre-existing table. On a fresh DB the CREATE TABLE carries the columns so
the indexes build; booting over the prior release's agent_sessions table
the CREATE TABLE IF NOT EXISTS no-ops and the index references a
not-yet-added column ("no such column: target"), failing the release
migration smoke. Build the two indexes after addColumns so the columns
exist first on both the fresh and upgrade paths.
TestMigrateOverLegacySessionsTable locks it over the pre-target schema.
hanzo code 402s ('a billable tenant is required') on a deployment without
IAM_MINT_CLIENT_*: codeToken() preferred the hk- API key, but an hk- key only
mints a billing principal when the server can resolve it (iamKeys.resolve is a
no-op without the mint credential) — so the request arrives anonymous and zen's
/v1/messages billing gate refuses it. The /v1/models catalog call is loosely
gated, so the launcher banner still printed, masking the cause.
Reorder codeToken to prefer a FRESH hanzo login JWT, which carries the caller's
owner/project/sub claims verbatim and mints a billing principal on EVERY
deployment, then fall back to the hk- key chain (still works on mint-credentialed
servers). HANZO_API_KEY stays the deliberate operator override at the top.
freshAccessToken() is a new expiry-guarded accessor (decomplected from
accessToken, which whoami keeps expiry-agnostic): an EXPIRED jwt must not win —
it would 401 a session a valid hk- key would serve — so it falls through to the
key. No expiry record (a raw HANZO_TOKEN) is trusted as-is.
Pinned by TestCodeTokenPrecedence (fresh jwt beats hk-, expired jwt falls
through, HANZO_API_KEY overrides all). Structural hanzoai/jwt extraction is a
follow-up PR.
The lifecycle reactors (notify / mirror-out / index-on-push) read the package
'mounted' var in DETACHED EmitLifecycle goroutines while Mount/Shutdown write it;
a reactor goroutine outliving a Shutdown (or a test Mount<->Shutdown cycle) tore
against the write. Convert 'mounted' to atomic.Pointer[cloud.Service[state]] —
Store on Mount/Shutdown, Load in every reader (reactors, export CloneURL/VerifyRef,
index-on-push, the GitHub importer) + the tests. Behavior unchanged (production
sets mounted once); the data race the -race detector flagged is gone.
Lands clients/connectorruntime — the native replacement for the standalone
ActivePieces Node engine: an ActivePieces JS connector action runs in goja
in-process (no auto pod, no in-cluster HTTP hop, no shared X-Piece-Run-Secret).
Same {action,auth,props} -> {ok,output,error} contract; org-gated via
principal.Org; the caller's credential travels in the request auth.
POST /v1/automations/connectors/:id/run run one connector action in-process
Wired into the composition root (subsystems.Wire) the one way main does it —
an explicit MountSpec after automations (it pairs with the catalogue) — not the
old self-registering cloud.Register init (removed; main is explicit-Wire).
frozen order sequence updated to 84 specs.
Re-derived from the cloud-cr worktree's uncommitted work against current main:
the package's principal.Tenant -> principal.Org, the cloud.Register init -> the
explicit Wire() entry, and the go.mod esbuild dep flipped to direct. The
kb/sync_piece.go refactor that originally accompanied this is NOT carried —
clients/kb was removed on main since the branch forked, so there is nothing to
refactor; connectorruntime stands on its own (imports only cloud + principal).
Build green; go test ./clients/connectorruntime/... green (registry + runtime);
TestWireOrderMatchesFrozen green at 84 specs.
Trace attribution, per-project metrics, sessions, and annotation-queue routes
on the o11y plane (clients/o11y annotation store + queues, clients/eval
attribution/telemetry, principal scope). Genuine new feature (not on main),
rebased clean onto current main; builds and race tests green.
Install the Hanzo GitHub App -> list org repos -> import into git.hanzo.ai ->
bidirectional sync (outbound mirror already exists; inbound = HMAC-verified
webhook, fast-forward-ONLY, never force-overwrites native). Inert until the App
creds land in KMS.
The monthly accrual latched once on the FIRST sweep's PARTIAL month-to-date spend
and no-oped every later sweep, so an affiliate whose dashboard swept early in the
month froze its share near zero (underpaid, though platform-safe). Store.Accrue now
inserts the period row on first sweep and, for the still-open period, TOPS IT UP
toward the current (higher) month-to-date reading in one transaction — adding only
the positive delta, so accrued_cents converges to the month-end value, never
decreases, and never overshoots. The per-event money invariant is untouched: each
row keeps margin + share from ONE reading (share <= margin) and the level-schedule
cap keeps sum(share) <= margin at every step.
Public link clicks move off the money-DB write path: clickLink folds pings into an
in-memory coalescing buffer (bounded), flushed batched on the next links read and on
shutdown, so a click flood can never contend with the accrual/payout writes.
Tests: TestAccrualConverges + TestAccrualConvergesAtMaxRate prove the share tracks a
growing month-to-date and sum(share) <= margin at every intermediate sweep; existing
invariant/idempotency/isolation/links tests stay green (26 pass, -race).
A push to a repo's default branch enqueues a durable IndexRepoWorkflow on
the embedded tasks engine (workflow id keyed by commit = idempotent per
push); a worker reads the tip tree from the object plane and folds its
text files into the org's code index, retried on failure. Fail-soft:
before the engine is wired the reactor indexes inline on its detached
lifecycle goroutine (the ai-ingest contract), so push-index is always live.
git and code never import each other — the reactor is git's third
lifecycle subscriber, the index reached through the SetIndexer func seam
wired once at the composition root.
eval telemetry (Trace): add ProjectID, SessionID, APIKeyHash (SHA-256 ref, never
plaintext), and StartTime/EndTime latency; DDL columns + additive migrations +
write/read paths. A run stamps its project (principal.ProjectScope), groups its
item-traces under the run as one session, times the model call, and records a
non-reversible ref of the caller credential. Trace list narrows by the caller's
project (default == whole org).
evals/metrics: thread the server-minted project scope into MetricsFilter,
usageWhere (AND project = ?) and the latency span filter. The default-project
(whole-org) board queries the ledger today; a named-project board is honest-empty
until cloud_usage carries a project column (ai write path) — activation is one
guard flip, the query plumbing is project-aware and tested.
principal.ProjectScope: the ONE helper for "default project == "" == whole org";
eval + o11y both read it.
o11y: explicit org-gated GET /v1/o11y/sessions pinning the runtime /api/sessions
route; native annotation-queues surface (SQLite metastore, org+project scoped) at
/v1/o11y/annotation-queues* — list/create/detail/update/delete, items add/list/
complete — returning the console {data,meta} envelope.
Tests: trace attribution (latency, session grouping, hashed-not-plaintext key),
per-project + cross-org trace isolation, usageWhere project predicate, annotation
queue lifecycle + org/project isolation + validation + principal gate.
integrations (github.go/github_app.go/github_webhook.go): App install ->
installation-token mint (ghinstallation), list granted repos, background import,
inbound push webhook (HMAC-verified, fast-forward-only). Inert until
GITHUB_APP_{ID,PRIVATE_KEY,WEBHOOK_SECRET,SLUG} land.
git (github_import.go + cloud.GitImporter seam in git_import.go): import =
fast-forward mirror-in per branch; inbound = fast-forward-ONLY advance where
native is canonical, so a divergence is a recorded conflict and never a
force-overwrite; loop-prevention via the Origin stamp; per-repo status. Outbound
mirror uses the org installation token for github.com targets, shared-token
fallback otherwise (no regression).
GET /v1/code/tree?repo= → get_repo_structure: the repo's files + per-file
symbol counts, ordered, per-org isolated.
GET /v1/code/file?repo=&path= → read_file: the INDEXED content (the symbol
chunks the search tiers hold), for fast context. NOT byte-verbatim — inter-
symbol lines a parser did not chunk are absent; the S3-backed git object
plane (clients/git) is the source of record for exact bytes, history, and
blame. Comments say so; a follow-up routes byte-fidelity read_file/blame at
the git plane.
Store.tree + Store.fileContent read the existing files/symbols/chunks tiers —
no schema change. Test covers tree structure, indexed-file read, 404 on an
unindexed path, and per-org isolation (org B sees an empty tree).
The coding-agent code lands fail-closed and INERT: no sandbox runs until an
operator provisions docker and sets HANZO_CODING_* env. The coding sandbox
stays DISABLED until the default-deny-egress + container resource-limit
hardening lands and red re-clears. Do NOT provision docker / set HANZO_CODING_*
until then.
Native coding tasks from Slack over the shared object plane: clients/coding
dispatcher, in-process agents session bus, tracker agent PR, bot NDJSON client,
git export seam, slack coding trigger.
Mission-control backend over /v1/agents/sessions: sessions carry host/cwd/
repo/target and a compact last-event; new /v1/agents/targets registry
(register/list/detail/patch/delete) with live session-load, org fail-closed,
in the same agents.db. Composes with the compute fleet at the view layer.
Sessions gain host/cwd/repo/target so mission-control can show where a
session runs and map sessions to a machine; the list projection carries a
compact last-event preview. Add /v1/agents/targets — register/list/detail/
patch/delete a dispatch destination (laptop|cloud|gpu|cluster|machine) with
a live session-load rollup — in the same agents.db, one tenancy column, org
fail-closed. A session's target resolves same-org at register/patch (#48).
Composes with the compute fleet (/v1/fleet/workers, /v1/clusters) at the
view layer rather than duplicating it.
Server-side copies alongside the ghcr retag; the cluster deploys from OUR
registry, ghcr stays the public identity. Best-effort by contract — a mirror
hiccup never blocks the tag-receipt.
Claude Code sizes a model's context window only from ids it recognizes.
An unknown gateway id like `zen5-pro` gets a hardcoded 128K budget and is
rejected client-side above it ("maximum context length is 131072") — even
though api.hanzo.ai and the DeepSeek-V4 upstream serve the full 1M window
fine (verified live: 140K-token prompts return 200).
Bridge each CC tier through a recognized carrier id (claude-opus-4-8[1m]
etc.) that unlocks the 1M budget, then rewrite it back to the served zen
alias via settings.json modelOverrides before the request leaves the
client — the carrier never reaches the server, so real claude-* routes are
untouched and responses stay branded/served as zen. One source of truth
(zenTiers) drives the wire env, picker branding, and overrides. The fast
tier stays direct (never needs >128K; SMALL_FAST_MODEL isn't override-
rewritten, so a carrier there would leak raw).
Verified end-to-end with Claude Code v2.1.210: the wire model reaching
api.hanzo.ai is zen5 / zen5-pro carrying context-1m-2025-08-07, with zero
raw claude-* leaks.
Claude-Session: https://claude.ai/code/session_01SpMZ69ur3tjAXCiwaa7Wv2
The store surface registers natively (api/store.Route on a group-scoped chain
mirroring the standalone /v1 bundle) instead of re-adding the gin AdaptNetHTTP
prefix; commercePrefixes keeps /v1/store pinned so store reads can never fall
through to the AI /v1/* balance gate again.
- subsystems/commerce.go: commerce.Embed(App: app) — the SharedApp contract;
commerce registers /v1/commerce/* + /_/commerce/* directly on cloud's router
(one specificity space, no second engine, no net/http adaptation)
- live wire paths registered natively with commerce's own gates:
/v1/billing/webhooks/:provider (HMAC-is-auth) and
/v1/billing/auto-recharge/run-all (service-token + PlatformOnly) —
commercePrefixes contract test pinned and passing
- commerceinproc.SetApp: S2S byte-stream enters the shared app's pipeline;
RoundTrip normalizes client-style RequestURI at the ONE seam;
entitlements + BalanceCents stay pure direct-Go (commerceclient)
- zip v1.8.2 (union tag: v1.8 features + the v1.7.4 chain-order fix + v1.7.5
empty-leaf normalization — v1.8.1 was missing both)
- metering dispatch e2e ported to the zip harness and green
The bare storefront surface api.Route(Group("/v1")) registers on the commerce
gin engine — GET /v1/store/current (the org-scoped default store the admin
dashboard AND the content storefront edge resolve), the /v1/store/:id/listing
upsert the publish edge writes, and the public /v1/store/:store/listing reads
karma.style serves at runtime — was dropped from commercePrefixes by the unfork.
Unowned, /v1/store/* fell through to the bare /v1/* AI catch-all, whose prepaid
LLM balance gate 402'd every store read for an org funded in commerce but $0 in
the ai ledger (org karma: $999.99 credit, GET /v1/store/current -> 402
insufficient_balance, blocking store provisioning + the storefront image fan-out).
Restore the /v1/store prefix so /v1/store/* routes to the commerce handler, which
resolves the org from the gateway X-Org-Id (TokenRequired -> ensureIAMOrg) and
lazily provisions the org's store — standalone-commerced parity. A store-metadata
read never requires an LLM balance. Per-route permission masks are unchanged
(money paths stay publishedRequired); gin still 404s an unknown /v1/store/* path.
Guarded by TestStoreSurfaceRoutedToCommerceNotAIGate (store read resolves on
commerce, never 402 at the catch-all) + the extended commercePrefixes pin test.
Add an optional media field (JSON array of URLs) to the social Post store,
additive and idempotent exactly like the marketing scheduled_at fix:
- store: media TEXT NOT NULL DEFAULT '[]' column + addColumn upgrade for
pre-existing prod DBs (migrate-on-open, the only upgrade path for the
encrypted single-file store); JSON encode/decode helpers; Post.Media []string
always serialized as [] never null.
- api: normMedia bounds each URL to maxField and the list to maxMedia (10),
the one sanitization seam on create + update (mirrors content/channel).
- tests: media round-trip in TestPostCRUD (create sets, update replaces) and
TestMigrateAddsMediaColumnToOldPosts — old-schema DB gains the column and a
media write no longer 500s; legacy rows default to [].
Margin-based accrual (share = rate x Hanzo's margin, share <= margin invariant),
derived share ledger that never mutates the cost-of-record, shareable links with
click/signup/conversion, per-period earnings, opt-in privacy leaderboard, and the
SuperAdmin set-rate. Tenant-isolated, fail-closed.
Switch the affiliate accrual from revenue-share (spend x rate) to PROFIT-share
(margin x rate), where margin = a referred org's gross spend x the platform
gross-margin fraction (AFFILIATE_MARGIN_BPS, default 40%). The share is a rate
OF Hanzo's margin, never the customer's bill, so a payout can never exceed the
margin earned. The L1 rate is capped (maxL1RateBps=9300) so the whole L1+L2+L3
schedule stays <= 100% of the margin: total share <= margin, always.
The derived share ledger (affiliate_accruals) now records the margin base
alongside the gross spend and the share; it never mutates the cost-of-record
(commerce/cloud_usage) - it is a pure projection keyed by referrer.
Dashboard surface (all org-scoped, fail-closed):
- GET /v1/affiliates/me/earnings per-period + per-direct-referral aggregate
- GET/POST /v1/affiliates/me/links shareable links + click/signup/conversion
- POST /v1/affiliates/me/handle opt-in leaderboard display name
- POST /v1/affiliates/click public click ping (vanity counter)
- GET /v1/affiliates/leaderboard opt-in handles + aggregate + your own rank
- POST /v1/admin/affiliates/:id/rate SuperAdmin set L1 rate (capped)
Tests: margin invariant (share <= margin per level + summed; charge unchanged),
set-rate cap, links lifecycle, cross-affiliate isolation (earnings + links),
leaderboard privacy (no org identity leaks, own rank always visible).
Address Red's LOW-1/LOW-3 on the production-header posture:
- serve.go sets zip.Config.ServerHeader=cfg.Brand so the responses the
ProductionHeaders middleware can't reach — the fasthttp transport's OWN
pre-routing errors (431/400) — read Server: <brand> instead of the framework
default. Requires zip>=v1.8.1 (transport now propagates ServerHeader). Repin
v1.8.0->v1.8.1; go.mod diff is only the zip line.
- BrandForHostOK strips a trailing FQDN dot so api.lux.network. resolves to lux
(fails safe to neutral before, never a wrong brand — brand-fidelity fix).
Wire the shared production response-header posture (zip v1.8.0) into the edge,
right after RequestID so it covers every response — success, error, 404, and
the public-site static bytes:
- Server is the white-label brand of the request Host via BrandForHostOK (cloud's
own registry), so a lux/zoo caller is never served "hanzo" and no response
leaks the framework name. An unmatched Host falls back to this deployment's own
brand (cfg.Brand), never a framework or hardcoded single brand.
- X-Api-Version carries the build version (Config.Version <- CLOUD_VERSION env,
else the link-time cloud.Version default) under a brand-neutral key.
- HSTS + nosniff are the always-safe security floor; the console SPA keeps its
own framing rules (no X-Frame-Options/CSP forced here).
Bumps zip v1.6.0 -> v1.8.0 for middleware.ProductionHeaders. X-Request-Id stays
owned by middleware.RequestID.
The Go hanzo code wired zero MCP — subagents got a bare model, no Hanzo tools.
Port the Rust CLI's resolve_mcp: resolve hanzo-mcp (installed → PATH, else uvx
hanzo-mcp), write an --mcp-config stdio document scoped to the cwd into the
isolated config dir, and pass --strict-mcp-config so the Hanzo server is the sole
MCP source (a repo .mcp.json is ignored — it can ship a bearer-exfiltrating stdio
server). A missing hanzo-mcp warns and continues (MCP is an enhancement, never a
blocker). So hanzo code claude now starts with the Hanzo tool lattice — code
search over the cloud index, web search, vision, fs/exec/git. Test covers the
opt-in, the config shape, and the not-found warn path.
Both indexers defaulted CLOUD_EMBED_MODEL to bge-m3 (the raw upstream), which
the gateway rejects (400) — only the zen-embedding SKU is served. Result:
index-time embeds failed silently, semantic code + KB search returned nothing
(vectors:0, degraded:true). Default to the served SKU. Pairs with the zen
bge-m3->zen-embedding alias so both the name and the SKU resolve.
Adds the gitlab provider on the same OAuth registry as slack/google/github.
Reads GITLAB_CLIENT_ID/GITLAB_CLIENT_SECRET from ENV (KMS-synced, never in code);
Configured() is false until both are present so it ships INERT and fails closed
(honest 503 / failure redirect, never a fake OK). Callback = the app's
/v1/integrations/gitlab/callback; the generic dispatcher seals the access+refresh
tokens into the org KMS namespace.
Least-privilege by construction: requests only openid/profile/email/read_api/
read_repository/write_repository — the token receives the intersection of
requested + app-allowed, so we never request api/sudo/admin_mode/k8s_proxy/
*_runner/*_registry even if the app was provisioned with them (tested).
GITLAB_URL supports self-hosted GitLab. 5 tests: registration, least-privilege
authorize URL, exchange seals both tokens + resolves account, missing-secret
fails honestly, error body surfaced.
commerce v1.47.2 fixes GET /v1/store/current returning the phantom shared
"default" store: it now resolves the caller org's namespace and lazily,
idempotently provisions the org-scoped store (store.EnsureDefault) on first
authenticated hit — the store id the content storefront edge needs to publish
Listing.headerImage. Round-trip + cross-tenant tests ship in the module.
Claude-Session: https://claude.ai/code/session_01Gq8suw7uuodAMPDRpo6iAB
POST/GET /v1/marketing/campaigns 500'd on prod with "table
marketing_campaigns has no column named scheduled_at": migrateCampaigns()
uses CREATE TABLE IF NOT EXISTS, which NEVER alters an existing table, so a
prod DB created before scheduled_at was added to the DDL was frozen at its
original schema and every campaign write (INSERT/UPDATE name it) 500'd. The
store is an encrypted single-file SQLite only the binary can open, so a
hand-patch is impossible — the upgrade MUST happen in migrate-on-open.
- migrateCampaigns: after the CREATE, run an idempotent additive column
upgrade — ALTER TABLE marketing_campaigns ADD COLUMN scheduled_at
INTEGER NOT NULL DEFAULT 0 — swallowing SQLite's "duplicate column name"
(the only error) so it's a no-op on a fresh DB.
- addColumn helper mirrors clients/social/store.go exactly (the ONE way we
do additive migrations), keyed by an extensible {table,col,def} list.
Tests (store_migrate_test.go, CGO_ENABLED=0 pure-Go cek): a from-OLD-schema DB
(marketing_campaigns without scheduled_at + a legacy row) opens, migrate ADDs
the column, a scheduled_at write succeeds, and the legacy row survives with
scheduled_at defaulted to 0; plus a fresh-DB idempotency re-open. Without the
ALTER the old-schema test reproduces the exact prod 500.
Claude-Session: https://claude.ai/code/session_01Gq8suw7uuodAMPDRpo6iAB
The unfork (6a071d2) rebuilt commercePrefixes in subsystems/commerce.go from a
pre-fix snapshot, dropping /v1/billing/auto-recharge (landed as #274/6dc3c6b)
and deleting its pin test with the old clients/commerce tree. Without the
prefix the durable cron's quarter-hour billing-autorecharge poke lands on the
account-bridge /v1/billing/* session gate and 403s — verified live before
6dc3c6b shipped in v1.801.1 (poke 200, 311 orgs swept, 23:45:05Z).
Same one-line-family fix, new canonical location; TestCommercePrefixesPinned
now lives beside the list it pins so a future rewrite can't silently regress
the wire path again.
Claude-Session: https://claude.ai/code/session_01XptqW83ZLpqyGBENc1wAQz
zenKeyResolver read the embedded KMS store ONLY. The operator injects the
provider keys (DO_AI_API_KEY, ANTHROPIC_API_KEY) as env from the KMS-synced
K8s secret cloud-api-llm-keys, but that value is not seeded into the embedded
ZapDB KMS store — so GetSecret missed, the resolver returned an empty key, and
zen's upstream call to DO GenAI answered 401 'Unable to authenticate you'.
Every zen chat failed while ai (which reads the key from env) worked.
Read env first, then KMS — the same order ai uses (object/kms.go). One key
source of truth shared by both zen and ai. Empty on both still returns '' so
the call fails fast, never silent free usage.
One canonical commerce repo, one-way dependency (cloud → commerce), no more
monorepo-split force-pushes to keep two trees in sync.
- clients/commerce/ (fork, no go.mod) DELETED; cloud imports the module
- subsystems/commerce.go: the ONE adapter — narrows cloud.Deps, boots
commerce.Embed, mounts the gin handler at the commerce prefixes, wires the
two in-process seams; luxfi/log imported plainly as log
- consumer bridges move cloud-side (they read cloud seams, not commerce):
clients/metering ← fork metering (finance-coupled billing-gate client)
clients/commerceclient ← in-process entitlement client + BalanceCents
(separate from commerceinproc: the entitlement client imports clients/plan,
which imports cloud — commerceinproc must stay stdlib-only for build.go)
- commerce API path: api/api flattened to api (module v1.47.1)
- middleware precedence regression tests moved INTO the module beside the
accesstoken fix they pin; the in-proc dispatch e2e stays in clients/metering
- subsystems/wire_test: freeze gitops (main had 83 specs vs 82 frozen)
Test surface green: subsystems, commerceclient (real embedded-ledger money
tests), commerceinproc, metering, catalogsync, bots, admin/finance, ml; root
package fails ONLY the 11 pre-existing env-gated tests (cek master key),
identical to origin/main.
Fable is Claude Code's top model tier; it was pinned to zen5-pro, the
same as Opus. Point it at zen5-max (the largest zen5 SKU: Qwen3.5-397B ->
1M overflow) so the CC tier ladder is monotonic: Haiku->zen5-flash,
Sonnet->zen5, Opus->zen5-pro, Fable->zen5-max.
Bumps the embedded o11y (v1.5.26->v1.5.28) and otel-collector
(v0.144.13->v1.2.0) to their koanf-v2 releases and drops the obsolete
otel-collector v0.144.10=>v0.144.13 replace. Removes the ambiguous
github.com/knadh/koanf/maps import (bundled v1.5.0 monolith vs split
module) that broke 'go build ./cmd/cloud'. Cloud builds green.
The ArgoCD-grade GitOps control plane over the operator App CRs, native to the
cloud binary and parallel to /v1/git. SuperAdmin-only, fail-closed, Secrets never
surfaced. The console dashboard consumes these shapes:
GET /v1/gitops/applications list: name, role, version(declared),
runningVersion, health, sync, phase, endpoints
GET /v1/gitops/{name}/tree flat node list (ArgoCD ApplicationTree) with
ownerRef parentRefs + per-node health
GET /v1/gitops/{name}/resource/{ref} live manifest + desired-vs-live diff
(ref = group:kind:namespace:name from a node)
GET /v1/gitops/{name}/logs newest app pod logs (tail/container bounded)
POST /v1/gitops/{name}/rollback pin CR image tag to a prior semver — REUSES
the P1 release seam (cloud.OnServiceRelease)
POST /v1/gitops/{name}/sync request an operator reconcile now
- health.go: per-resource health in the ArgoCD vocabulary (Healthy/Progressing/
Degraded/Suspended/Missing), pure — P2b swaps to gitops-engine pkg/health.
- CR kind-collapse compat shim: reads BOTH apps.hanzo.ai (kind App, forward) and
services.hanzo.ai (kind Service, live) — App wins the dedupe; removable
post-cutover. spec.role surfaced.
- Wired into subsystems.Wire after paas (so the release seam is registered before a
rollback delegates to it).
TODO seam (follow-on, noted): true GitOps on git.hanzo.ai — RegisterPushBuilder
commits the CR change to the manifest repo and the engine syncs repo→cluster;
desiredSource flips last-applied → git with no shape change.
Tests: health matrix, sync, ref-parse, membership, diff, observe, App-first/Service
-fallback resolution, list dedupe, tree ownerRef+selector. go build + test green.
The two-boot migration smoke shares a docker named volume across boots, but a fresh
named volume is root:root 0755 while the cloud image runs non-root — so the baseline
(v1.799.19) could not create cek's <db>.cek.lock under /data and aborted before
"listening" ("cek: open lock ... permission denied"), failing the gate on infra, not a
regression. The single-boot smoke only passed because --tmpfs is world-writable.
chmod the shared volume 0777 via a root helper before the baseline boot, and again
between boots so the candidate can read the baseline's files even if runtime UIDs differ.
The durable platform cron (clients/cron) fires cron-billing-autorecharge every
15m as a poke: POST cloud.hanzo.svc:8000/v1/billing/auto-recharge/run-all with
the COMMERCE_SERVICE_TOKEN bearer. That path was not a commerce prefix, so it
fell through to the account-bridge /v1/billing/* catch-all, whose session gate
403s a service token ("sign in to view billing"). Live fires have been failing
on exactly that — verified in-pod: the poke returns 403 on v1.799.13.
Add /v1/billing/auto-recharge to commercePrefixes alongside /v1/billing/webhooks
— same class of route (token/signature IS the auth, no session possible).
commerce mounts at Wire order 100, ahead of the bridge, so the poke reaches gin
where commerce's own TokenRequired service-token branch + PlatformOnly gate
authenticate it. No other /v1/billing/* route changes owner.
TestAutoRechargePrefixMounted pins both prefixes so a future edit can't silently
re-break the sweep. Landed directly on main: the identical change merged four
times (#274/#275/#277/#280) and was each time force-pushed off main or closed +
branch-deleted; a PR is not a durable landing surface here.
Claude-Session: https://claude.ai/code/session_01XptqW83ZLpqyGBENc1wAQz
tracker.migrate() indexed issues(org, repo) and issues(org, kind) in the base
DDL, but repo/kind are ALTER-added. On a legacy tracker.db (CREATE TABLE IF NOT
EXISTS no-ops) those indexes fail "no such column", migrate() fails, mount
fails, and the pod crashloops on deploy — the same class already fixed in
wallets and affiliates. Move both indexes after the ALTER pass.
Prevent recurrence with a shared regression harness (internal/migratetest):
each store contributes a legacy-DDL case that seeds its pre-migration schema and
asserts migrate() succeeds and is idempotent, proving migration-correctness as a
pure test over schema epochs instead of at prod boot. Cases added for tracker
(with a scoped-insert probe) plus the other ALTER+index stores — agents, social,
platform, provisioning, projects — which audit clean and are now locked.
The plain smoke boots on a fresh /data, so every migrate() takes its CREATE-TABLE
path and no forward-migration runs — an index over a not-yet-ADDed column is valid
on a fresh store yet crashes on a pre-existing one (affiliates referrer_org in
v1.800.1, wallets project/agent before it), which is how a boot-crash reached prod
and took api.hanzo.ai down while every smoke stayed green.
Reproduce the real upgrade path: boot the prior released image (default
ghcr.io/hanzoai/cloud:v1.799.19, override via SMOKE_MIGRATION_BASELINE) to lay its
cek-encrypted on-disk schema into a persistent volume under one shared throwaway
master key, then boot the candidate over the same volume and require "listening". A
migrate() that assumes a fresh store fails the gate before any image is pushed.
Close push→build→image→CR: a proven, clean-semver image rolls live by patching
the matching operator hanzo.ai/v1 Service CR's spec.image directly, so the
operator reconciles the Deployment. This is the in-cluster, direct-CR replacement
for universe's image-update.yml GitOps hop (repository_dispatch → PR → ArgoCD),
with the same determinism (clean-semver only; resolve CR by metadata.name) and no
git round-trip.
- build.go: RegisterServiceReleaser / OnServiceRelease / ServiceReleaserRegistered
inversion seam (mirrors RegisterPushBuilder), so a build-completion path rolls a
proven image with no cloud⇄paas import cycle.
- clients/paas/release.go: releaseService — resolve CR by name (main-first),
clean-semver gate (IsSemverTag), idempotent merge-patch of spec.image; registered
at Mount as the releaser impl (paas owns the first-party Service CR plane).
- clients/platform/release.go: rolloutRelease — native CR patch primary, universe
image-update dispatch kept as an additive GitOps mirror during cutover.
Tests: split/gate, patch, idempotent, reject-floating, unknown-service, main-first,
fail-closed, seam dispatch/no-op. go build + go test green (paas, platform, root).
On a store whose affiliate_referrals table predates referrer_org, CREATE TABLE
IF NOT EXISTS is a no-op, so 'CREATE INDEX ... ON affiliate_referrals(referrer_org)'
in the same DDL batch failed with 'no such column: referrer_org' BEFORE the
ALTER ... ADD COLUMN pass ran — crashing cloud on boot (v1.800.1 CrashLoopBackOff,
api.hanzo.ai down). Move the index creation after ADD COLUMN so it is valid on
both a fresh store and a migrated one. Regression test seeds the old schema and
asserts migrate succeeds + backfills (it fails with the exact prod error when the
index is moved back into the DDL batch).
`hanzo code claude` already pins the model to a zen5 alias (default zen5, the
GLM-5.2-class tier) and forces it on argv so a persisted /model selection
("best") cannot override it — but Claude Code's base system prompt still tells
the model it is Claude, so a Hanzo-served model self-identifies as Claude when
asked. Append the Hanzo Zen identity via --append-system-prompt (an APPEND, not
--system-prompt: CC keeps its harness prompt for tool-use/safety/coding) so the
served model says it is a Hanzo Zen model. The identity is not a permission
bypass, so it is applied in --safe too (unlike --dangerously-skip-permissions).
codex/dev (OpenAI wire) are unchanged — the append is Anthropic-only.
- `hanzo` (no args): log in if needed, then drop into the configured agent on
a Hanzo cloud model — one word, billed to your account, all zen-native.
- `hanzo code` (no agent): runs the default agent instead of showing help.
- Default agent resolves HANZO_CODE_TOOL, then config `code_tool`, else dev.
- `hanzo config set code_tool claude|codex|dev` + `code_model` — git-style k/v,
persisted to ~/.hanzo/config.
A named agent (`hanzo code claude`) still dispatches to its subcommand.
migrate() created ix_wallets_scope/ix_wallets_finance over project/agent/
finance_account inside the base DDL, before the idempotent ALTER TABLE that
forward-adds those columns. On a wallets table created before scoping (the prod
shape), CREATE TABLE IF NOT EXISTS no-ops, so the index build hit 'no such
column: project' -> migrate fails -> mount fails -> pod crashloop on any newer
image. Order by dependency: base tables, then ALTER-add every post-original
column (project/agent/chain/finance_account), then the indexes over them.
Regression test opens a legacy-schema wallets.db and asserts clean, idempotent
migrate + a fully-scoped insert.
The identity boundary validated JWTs only; an opaque API key (hk-/sk-/pk-)
yielded no principal, so a subsystem that gates on the minted identity (zen's
billing gate) refused a key request as anonymous — the zen 402 'a billable
tenant is required' for a funded hk- key.
keyResolver turns a key into the SAME idClaims a JWT yields (via IAM's
authenticated get-user?accessKey, the confidential hanzo-console client), so the
ONE minting path serves both credentials and key auth and session auth can never
disagree on who a request is. An unresolved key stays anonymous — a bad key
never grants trust; an unconfigured resolver keeps keys anonymous rather than
mis-resolved. Brief generic TTL cache keeps the hot path off the network.
Red LOW-6: a timed-out coding run must still close the session and mirror
its terminal result. Run terminal-side ops (fail/CloseSession/mirror/CreatePR)
under context.WithoutCancel so an expired run ctx cannot strand a session in
'running'.
Red MEDIUM-3: the cloud→bot coding POST carries the org hk- git credential and
the shared gateway bearer. Fail closed on a cleartext http:// target; a plaintext
in-cluster hop is allowed ONLY when the operator asserts mesh mTLS via
BOT_GATEWAY_ALLOW_PLAINTEXT=1. Error never echoes the credential.
Turn @hanzo from a chatbot into an engineer. A Slack message
`@hanzo code: <repo> <task>` branches off the chat-only reply into a
durable coding run: register a live agent session, dispatch to the
bot-gateway sandbox, mirror progress into the session live, verify the
pushed branch landed in native /v1/git, open a native PR work item, and
report the branch + PR back in-thread. Non-code mentions keep the
existing chat path unchanged.
- clients/coding: transport-agnostic orchestrator (Dispatcher over
interface seams; unit-tested with fakes, org-isolated, no credential
leak into session events/PR body). Cannot import git (cycle via
integrations) so CloneURL/VerifyRef are injected at the composition root.
- clients/agents/inproc: in-process session API (Open/Log/Close) — the
twin of the /v1/agents/sessions control plane, same store + live bus.
- clients/tracker/agentpr: in-process Kind:pr Source:agent work item,
org-scoped, get-or-create repo board (KEY-N).
- clients/bot/coding: in-process NDJSON client for POST /v1/coding-tasks;
credential travels in the body only, never argv/URL/logs.
- clients/git/export: CloneURL + VerifyRef seams (org-scoped ref check).
- clients/integrations/slack_coding: the code: trigger, credential fetch
from KMS (fail-closed), ack, detached bounded run, result Block Kit card.
- subsystems/wire_seams: compose the Dispatcher (git seams + adapters)
and inject into the Slack surface.
Tenant isolation fail-closed: org is the only tenant key on every seam;
sandbox pointed only at the caller org's clone URL with an IAM-scoped
credential; cross-org repo targeting is refused by git's path-vs-identity
guard. Tests: CGO_ENABLED=0 go test green across all touched packages.
Pulls in the reworked zen family (one SKU per capability, every upstream
verified on DO) and the hanzoai/thinking depth fold (effort → each upstream's
native reasoning shape). Adds hanzoai/thinking v0.1.0 as a direct require.
zen5 is the flagship GLM-5.2-class alias (1M ctx, tool-capable — the glm-5.2
upstream returns tool_use/stop_reason:tool_use, verified live). Matches the
intent to launch Claude Code on GLM-5.2 through api.hanzo.ai. Was zen5-pro
(DeepSeek). The four CC tier slots stay pinned to served zen5 aliases so the
classifier, subagents, and /compact never hit an unserved claude-* id.
ev.Branch was an un-escaped mrkdwn sink: git refnames allow < > & !, and the
receive-pack fire path (branchTips → for-each-ref) applies no branchRE, so a
hostile branch (e.g. x<!channel>y) flowed verbatim into the summary and the
*Branch* field of every subscribed channel. slackEscape it in both places, and
defensively escape the org/repo display text too (a deploy event's repo derives
from an unconstrained RepoURL, though it is subscription-gated to a valid name).
TestNotifyEscapesMrkdwn now also pushes a hostile branch through smart-HTTP with
the real git CLI (go-git rejects such a refspec) and asserts it is neutralized in
both the summary and the Branch field.
HIGH-1 mirror-out no longer starves the shared pack plane: dedicated mirrorSem
(separate from packSem), per-push context.WithTimeout, git http.lowSpeed abort,
and cmd.WaitDelay so a stalled downstream's network-helper child can't wedge the
slot past the deadline.
MED-2/MED-3 full (org,project,repo) identity: subscriptions + mirror targets key
on project too (DDL + every list/delete + the notify + mirror fire paths); deploy
emitters thread project (platform a.ProjectID normalized, projects org-level);
repo delete cascade-deletes its subscriptions + mirrors in one tx (no
exfil-on-recreate via an orphaned target).
MED-1 decouple allowlists: outbound mirror TARGET set = {github.com, gitlab.com}
only; the local git host is rejected as a target (no internal SSRF /
privileged-cred presentation). Inbound-fetch credential gate unchanged.
MED-4 Slack mrkdwn escaping of user-derived text (commit subject, pusher, deploy
detail) so a crafted commit subject can't inject <!channel>/disguised links.
LOW-1 keep the shared bot token least-privilege (no chat:write.public — it would
also arm the @hanzo assistant to post uninvited); notifications require the bot
be invited. INFO: reject non-deliverable build.started at subscribe time; DRY
repoFromURL into cloud.RepoFromCloneURL.
Re-verified: new tests for project-scoped routing, delete-cascade, mrkdwn escape,
and stalled-downstream isolation; suites green under -race; gofmt+vet clean;
go.mod untouched.
* analytics: add capture (write) plane — POST /v1/analytics + /v1/tracker → hanzo.events
The analytics subsystem served only read lenses over hanzo.events; nothing
wrote the table, so the web/commerce lenses were permanently honest-empty. This
adds the symmetric ingest: products POST batches to cloud (the ONE native front
door) and cloud writes org-scoped rows into the datastore warehouse the read
side already queries.
- POST /v1/analytics, /v1/analytics/batch, /v1/tracker (beacon alias) — all
tenant-gated in-handler; tenant_id is always principal.Org, never client input.
- Writes ride ai/object.DatastoreExec (the SAME pooled client the reads use).
- The writer owns the hanzo.events DDL (EnsureEventsTable, idempotent/latched).
- Privacy scrub: credential/PII-shaped property keys dropped, email values
redacted, before any row is built.
- Pure core (normalizeEvent/scrubProps/buildEventsInsert) unit-tested; HTTP
contract tests cover no-principal 403, forged-org 403, oversized 400,
datastore-down 503; a build-tagged live test proves the full round trip
against a real datastore.
* analytics: accept anonymous capture, attributed to the brand-public org
Marketing sites emit anonymous pageviews (no session). captureTenant now falls
back — when there is no validated principal — to the PUBLIC brand org derived
SERVER-SIDE from the request Host via the white-label registry (BrandForHostOK),
never a client-claimed org. A forged X-Org-Id is still ignored, and an
unrecognized Host is refused (anonymous events are never dumped into a default
org). Gated by CLOUD_ANALYTICS_PUBLIC_CAPTURE (default on, matching the existing
public insights-capture posture). Verified live: an anonymous pageview to
Host hanzo.ai lands under tenant_id=hanzo.
---------
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Fold the remaining GTM subsystems into clients/marketing (native Go, per-org
SQLite, twin of clients/crm), so nothing Python is load-bearing in the mount
path:
- Email drip sequences on the embedded hanzoai/tasks engine (drip.go): a
per-minute durable schedule sweeps due enrollments; each (enrollment, step)
is claimed once so a step delivers at most once across restarts/redelivery,
then the walk advances or completes. Mirrors clients/cron (engine owns time,
SQLite owns the schedule) — no Redis, no bespoke ticker.
- The ONE send seam (suppress.go): every marketing delivery funnels through
state.deliver, which enforces the per-org suppression/opt-out list and then
hands off to the platform notify rail. Expose notify.Send so the existing
sender is reachable in-process — one sender, not a second. Plus a signed
public one-click unsubscribe.
- Audiences (audiences.go): cohort filters evaluated live against the org's
hanzo.events analytics via the ai/object datastore, tenant_id-scoped and
honest-empty when the warehouse is not wired.
- Promo codes (promos.go): the First-1,000 90%-off launch promo (discounts.md)
realized as a non-cash wallet credit through the finance ledger, with the
hard 1,000 cap, one-per-org, one-per-instrument and team-seat-cap guards.
- Content calendar (calendar.go): scheduled posts as documents published by a
task-executed hook; social publish returns an honest 501 until a connector
is wired (clients/social's push is fail-closed, the automations connector
registry is package-private).
- Campaign scheduling (scheduled_at + the scheduled state).
Real tests: drip scheduling + per-step idempotence + tenant isolation,
suppression enforcement at the gate, promo eligibility math + abuse guards.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Generalize the single-registrant git push→deploy hook into a many-subscriber
lifecycle stream without regressing push→deploy. RegisterPushBuilder/OnGitPush
stay exactly as-is (deploy is the subscriber-of-record); alongside them a new
RegisterLifecycleSubscriber/EmitLifecycle fans ONE LifecycleEvent out to N
reactors, best-effort, detached, on a cancel-immune context, panic-contained.
Emit points: PushLanded at the one branch-build funnel (covers HTTP/SSH/push,
now carrying before/after tips + pusher); BuildStarted/DeployLive/DeployFailed
at the platform (startGitBuild / applyLive / failDeploymentCtx) and projects
(deployGit / deployArtifact / completeDeployment) transitions.
Subscriber 1 — Slack notify: per-org repo→channel subscriptions
(/v1/git/repos/:name/subscriptions) stored in the existing per-org git.db,
delivered as Block Kit via the ONE integrations chat.postMessage path
(integrations.NotifySlack/PostSlackBlocks; the automations connector now shares
it). Adds chat:write.public so a freshly-subscribed public channel works
without a manual invite.
Subscriber 2 — outbound mirror: per-repo downstream targets
(/v1/git/repos/:name/mirrors) force-push ONLY the advanced branch to
allowlisted hosts (github.com/gitlab.com/git.hanzo.ai), token via env-only
http.extraHeader under the pack-slot semaphore. LifecycleEvent.Origin is the
loop-prevention seam for a future inbound sync.
All routes org-scoped + fail-closed; tenant isolation, injection, allowlist,
and loop-prevention covered by tests.
Upgrade the in-process Base embed from the single-instance waitlist-only fold
(#193/#211/#248) into two orthogonal lanes on the ONE engine:
- LANE 1 (unchanged surface): the platform waitlist app → public /v1/waitlist/*.
- LANE 2 (new): managed Base hosting → ONE Base app PER ORG, opened lazily and
LRU-pooled, each on its own SQLite under {DataDir}/base/{TenantSegment}/ (the
HIP-0302 'SQLite per tenant' model the gojabase leaves use). Served
authenticated under /v1/base/*, the org resolved from the VALIDATED cloud
principal (principal.Org) — physical per-org isolation, the console Bases
manager's backend and the in-binary replacement for the superbase pod.
Base serves under BASE_API_PREFIX=/v1/base so its collections API mounts natively
(self-URLs included) without colliding with cloud's other /v1 routes; the waitlist
plugin binds a FIXED /v1/waitlist regardless. Per-org apps validate bearers against
Hanzo IAM's JWKS as their exclusive auth source (apis.StoreKey{JWKSURL,
ExternalAuthOnly}) — ONE IAM, no second auth path. The cloud binary owns the ZAP
transport, so embedded apps set ZAP_DISABLED. Shutdown releases the platform app +
every pooled per-org app.
Reuses gojabase.TenantSegment (the ONE injective, traversal-safe org->path encoder).
Test: subsystem boots in the harness, a collection record round-trips per-org
isolated over the real HTTP path (acme's record invisible to globex; distinct
on-disk dirs).
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Adds clients/company (/v1/company): one formation state machine per org
(structure → founders+KYC → $999 → documents → esign → on-chain equity
genesis → company), with a SKIP path for already-incorporated orgs that
imports corporate docs → data room and a cap-table sheet → captable.
- machine.go: pure transition table + per-edge guards; the payment gate,
KYC gate, and skip path are unit-tested with no I/O.
- providers.go: narrow seams for billing/kyc/docs/esign/captable/anchor/
filing/upgrade. Billing wires the shared ResourceMeter ($999). KYC and
state filing are honest stubs (no fabricated verification/filing).
- genesis.go: KMS-signed Hanzo-L1 equity-genesis anchor mirroring
clients/treasury; computes the root always, commits on-chain when wired,
honest pending otherwise.
- adapters.go + captable.facade + dataroom.Ingest: new in-proc facades so
company writes the cap table and data room without an HTTP hop.
- google provider completed in clients/integrations (OAuth + KMS token
custody); the automations google_sheets/google_drive connectors are
consolidated to one `google` connector sharing that token; company import
reads Drive/Sheets through it.
- wired into subsystems.Wire() after referrals; docs/company-dogfood.md
walks Hanzo/Lux/Zoo through the import path.
clients/guide + /v1/guide/*: a per-org checklist engine over a
machine-readable curriculum (Step: id/title/why/how/done/dependencies/
signal/tool). Per-org progress on cloud.OrgStore; next-step + dependency
gating are pure functions; auto-detect reconciles a step to done when its
signal maps to real org state (acted = agent action ledger; analytics =
shared warehouse events). Business AI 'do it for me' drafts with deps.AI
then executes the step's bound MCP tool via automations.InvokeTool AS THE
CALLER — attributable, metered, audited, never exceeding the caller's
authorization. Built-in default.yaml (7 steps: positioning→landing→
analytics→waitlist→email→referral→launch) seeds it before marketing's
checklist.yaml; an org-custom PUT replaces it cleanly.
clients/automations: decomplect tool dispatch (dispatchTool) from its two
doors — the HTTP MCP handler and the new in-process InvokeTool/ToolExists
seam — so both share one dispatch, org-scoped credential, concurrency
bound, meter and audit.
Tests (pure-Go, CGO_ENABLED=0): validation/cycle-detection, next-step +
dependency gating, auto-detect reconcile (present/absent/error/terminal),
the acted detector end-to-end, the agent (tool success/failure/assisted/
unknown-tool), and the HTTP surface (403 gate, transitions, 409 gating,
curriculum replace/revert, per-org isolation, do-delegation).
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Extend the existing affiliate accrual into a depth-capped multi-level upline and
fold OSS-author royalties into the SAME per-org spend walk, over one attribution
spine.
Affiliates — multi-level upline (clients/affiliates):
- The referredBy graph is the existing affiliate_referrals edge made walkable
(referred_org -> referrer_org, denormalized). referredBy is set-once/immutable
(UNIQUE referred_org) with cycle detection at set time (wouldCycleOrg climbs the
proposed referrer's upline and refuses a loop).
- Per-level schedule L1 20% / L2 5% / L3 2%, depth cap 3. L1 uses the affiliate's
own negotiable rate (default 20%, preserving prior single-level behavior); L2/L3
are platform constants. accrual rows carry the level for analytics.
- The admin sweep is source-centric: fold over every referred org, read spend ONCE,
walk the upline and accrue to each ancestor's approved affiliate, latched
at-most-once per (affiliate, source, period). The per-affiliate dashboard read
walks the downline (same latch key, never double-accrues).
- User-level referredBy graph (user_referrals): set-once + cycle-checked, mirroring
the org edge; recorded from the referee's user to the affiliate's owner user.
- Payout no-overdraw guarantee unchanged (pending-guard + treasury reserve backing).
Authors — OSS royalties (clients/authors):
- Default author share 25% (was 5%).
- GitLab provider alongside GitHub: host-aware repo canonicalization (github.com +
gitlab.com), provider-dispatched forge seam (linkedAccount/repoAdmin/fetchFile).
- Append-only, on-chain-ready royalty ledger (author_ledger) with a nullable
compute_proof column, written per accrual in the same transaction as the balance
move. compute_proof stays NULL — the hanzod attestation is a follow-up, not faked.
- AccrueForOrg seam: the affiliate sweep drives author royalty for each source org
with the spend it already read — one accrual walk. Nil-safe when authors is
unmounted (mirrors treasury.Reserve).
Surfaces:
- GET /v1/affiliates/me — my code, link, downline by level (L1/L2/L3 with rates +
counts), accrued/pending/paid, payouts.
- GET /v1/admin/referrals — unified SuperAdmin cross-tenant analytics (top referrers,
conversion, accrual liability by level). The one-time-bonus board moves to
GET /v1/admin/referrals/bonuses (referrals) so the two compose without colliding.
Tests: 3-level walk math + depth cap, cycle rejection (org + user), set-once
immutability (org + user), no-overdraw payout, GitLab verify + ledger row with NULL
compute-proof, both surfaces. Deterministic fiber test timeout (30s ceiling; the 1s
default flaked under machine load).
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Complete the native knowledge graph on the kb subsystem.
Wikilink edges: the kb-page after_save hook now extracts [[Page Title]]
references from the page body (the same flattened Lexical text the vector
indexer embeds) and reconciles them into kb-link edge documents — a Link
(source) + Data (target_title) reference, no parallel store. Index and link
maintenance run as one combined page hook so a vector outage never skips link
extraction. Targets resolve to a page by value at read time, so a rename or
trash of a target needs no edge rewrite; a page's own trash removes its
outgoing edges.
GET /v1/kb/graph: the org's knowledge as nodes (kb-page/kb-memory/kb-source,
plus connector and dangling-link endpoints) and edges (parent tree, wikilinks,
connector provenance), org/project scoped, shaped for a force-directed
renderer.
POST /v1/kb/import: an Obsidian-importer-equivalent that ingests an Obsidian
vault zip, Notion export zip (markdown/HTML), Evernote .enex, or Roam JSON as
a kb-page tree with links preserved. Each format is a pure normalizer package
(obsidian/notion/roam/evernote over vault + lexical), mapping to pages filed
through the same framework.Ingest path a connector sync uses — the after_save
hook then extracts their wikilinks, one link path for authored and imported
pages alike.
framework: add the in-process Delete (twin of the HTTP delete, runs on_trash)
used by edge reconciliation.
Tests: table-driven wikilink extraction; per-format normalizer tests with real
fixture files; end-to-end graph, import, and full edge lifecycle (extract,
reconcile-on-edit, source-trash cleanup, target-trash dangling) over the real
framework store.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Wallets custody scoping (clients/wallets)
- One Scope type {org, project, agent, account} is the ONE key both the KMS
secret ref (keyRef) and the store lookup derive from. Org stays the hard
isolation boundary; project/agent/account are optional narrowings within it.
- keyRef derives from the full scope, injection-safe (narrowings validated to a
slash-free segment). An org-only wallet keeps its exact legacy ref, so scoping
is additive, not a migration.
- listWalletsByScope is the one scope-filtered read path (org bound, narrowings
filter within); create + list handlers thread the scope. Store gains project +
agent columns (forward ALTER, dup-column tolerant).
Native x402 pay-per-use subsystem (clients/x402, /v1/x402)
- challenge (402 + PaymentRequirements) -> client signs ERC-3009 -> proof ->
Verify (EIP-712 secp256k1 recovery via luxfi/crypto, the same primitive wallets
signs with) -> settle -> serve. Enforce is a zip middleware a priced route
group applies.
- Idempotent + replay-safe: settlement id is deterministic in (from, nonce); a
spent nonce reused for different terms is a replay (402), a re-submitted
authorization is an idempotent retry (settled + metered once). Two independent
guards: the PK-atomic store claim and the ledger's own RequestID/Ref idempotency.
- Settlement wires into the metering spine: the payer's org is debited through
metering so paid usage appears in billing/usage like any metered spend, and the
recipient wallet's ledger is credited. Ledger settlement is LIVE; on-chain
broadcast of the authorization is a seam (not wired).
- Marketplace seam: a Registry (Publish) maps resource -> Terms (price + recipient
wallet ref); x402 resolves the recipient via wallets.ResolvePaymentTarget and
enforces. The registry itself is another subsystem's work.
Removes the dead, unreferenced clients/commerce/payment/x402 (gin + btcec) so the
binary has exactly one x402.
Tests: scope derivation + injection + scoped seal/sign + scope lookup isolation;
EIP-712 verify round-trip/tamper/term-binding; challenge->verify->serve, nonce
replay rejection, settle-once on retry, free passthrough, payer-required, and
end-to-end ledger settlement (payer debited once, recipient credited once).
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
A repo populated by fetch/receive-pack accumulates objects via index-pack,
which writes a pack but no reachability bitmap and no commit-graph — so every
upload-pack walks the whole object graph to build the clone pack (O(objects)
per clone). Add maintenance on the git-exec seam: repack -adb --write-bitmap-index
+ commit-graph write --reachable --changed-paths, under one pack slot with the
same memory bounds as every pack op (a multi-GB gc can't OOM the pod).
- POST /v1/git/repos/:name/gc — on-demand repack, org-scoped, 404 fail-closed.
- Post-push autoMaintain: fire-and-forget git gc --auto after every receive-pack
(HTTP + SSH), slot-yielding (skips when clones are busy) so repos self-maintain
as pushes accumulate packs, no scheduler.
Tests: runMaintenance writes bitmap + commit-graph and the repo stays clonable;
the /gc endpoint repacks (200 maintained) + 404s unknown/cross-tenant repos.
Fiber's Body() transparently inflates a Content-Encoding: gzip request per
its header (BodyGunzipWithLimit + SetBodyRaw), so c.Body() is already the
decoded pkt-line stream. packRequestBody inflated it a second time and
returned 400 "invalid gzip request body" for every compressed clone/fetch.
git gzips the ref-negotiation once a repo carries enough refs, so this broke
clone for all but trivially small repos — the small-repo round-trip tests
sent the request plain and never exercised the compressed path.
Read c.Body() verbatim (the framework owns Content-Encoding). Add a gzipped
upload-pack regression test through the live Fiber server that fails on the
double-decode and passes on the fix.
Ships the zen consolidation: ai discovers the Zen family from the zen service
(GET $ZEN_URL/v1/models) and fronts it, holding no zen routing/pricing/identity
of its own. Pulls the corrected per-upstream reasoning vocabulary and the
money/decimal billing types transitively.
Address the red-team findings on feat/git-streaming-object-plane:
HIGH-1 mirror credential exfiltration: only attach GIT_MIRROR_TOKEN to hosts
on an allowlist (GIT_MIRROR_ALLOW_HOSTS, default {github.com, git.hanzo.ai});
any other https source fetches anonymously so a tenant-supplied URL can't
capture the shared token. Set http.followRedirects=false on the mirror fetch
so the token can't ride a cross-host redirect.
MED-2 pack RAM: add -c pack.threads=1 -c pack.windowMemory=64m
-c pack.deltaCacheSize=64m -c core.bigFileThreshold=16m to every pack/fetch
subprocess, and a concurrency semaphore (GIT_PACK_MAX_CONCURRENCY, default 2)
around all pack ops so concurrent large clones/mirrors can't multiply RAM in
the shared cgroup.
MED-3 org traversal: gate the org identifier through orgRE (alnum-led, no
'/'/'\\', no leading '.') in org(), so a SuperAdmin X-Org-Id switch to
"../../etc" is rejected at the git boundary before it reaches absRepoPath.
MED-4 mirror SSRF: resolve the source host and refuse loopback / private /
link-local / metadata (IMDS) / unspecified / multicast targets;
GIT_MIRROR_ALLOW_PRIVATE_HOSTS allowlists internal hosts for tests /
deliberate internal mirrors. Generic rejection message (no probe oracle).
MED-5 pack-to-disk DoS: -c receive.maxInputSize (GIT_RECEIVE_MAX_INPUT_SIZE,
default 2g) on receive-pack so a gzip-amplified or runaway push can't fill
the pod disk.
MED-6 disconnect reaping: gitPackStream.Close now closes the read pipe and
Kills the process before Wait, and runPackSSH Kills on a channel-copy error,
so an abandoned clone can't leave git blocked on a full pipe (leaked
proc/goroutine/FDs).
LOW-7 strip URL userinfo in mirrorSource (credentials via env only, never a
ps-visible argv). LOW-8 close std pipes on cmd.Start failure.
Tests: org-traversal rejection, disconnect-reaping (Close returns promptly +
process reaped + slot released), mirror credential host allowlist, and mirror
SSRF + userinfo strip; existing suites stay green.
The heavy git paths buffered whole packs in RAM via go-git's pure-Go
server transport and FetchContext: a clone serialized the entire outgoing
packfile into a bytes.Buffer, a mirror indexed the whole incoming pack in
memory, and receive-pack read the full push body. Mirroring a multi-GB
repo OOM-killed the 1 Gi cloud pod.
Route the object plane through the streaming git CLI instead — the way
gitea/GitLab serve smart-HTTP — so packs stream to and from disk with
memory bounded by an OS pipe:
- clone/fetch serve: `git upload-pack --stateless-rpc`, git stdout streamed
straight to the HTTP response (SendStream); no pack buffer.
- push receive: `git receive-pack --stateless-rpc`, request body -> git
stdin -> index-pack to disk; only the small report-status is buffered.
- info/refs: `git <svc> --stateless-rpc --advertise-refs` + pkt-line header.
- mirror-in: `git fetch --prune --tags +refs/*:refs/*` against the on-disk
bare repo; ls-remote --symref resolves the source default branch for HEAD.
- SSH: plain `git upload-pack`/`git receive-pack` over the channel.
One git-exec seam (gitexec.go) builds every subprocess with a hardened,
minimal env (GIT_CONFIG_NOSYSTEM, GIT_CONFIG_GLOBAL=/dev/null, no inherited
secrets, GIT_TERMINAL_PROMPT=0, GIT_NO_REPLACE_OBJECTS) and arg slices only.
Protocol v2 is forwarded from the client Git-Protocol header / SSH env
(validated). Mirror source credentials are injected only via env git-config
http.extraHeader (GIT_MIRROR_TOKEN), never argv or logs, under
GIT_ALLOW_PROTOCOL=http:https. Tenant isolation stays on the validated
absolute bare-repo path (storage.absRepoPath); the handler org/path guards
are unchanged.
Push-to-deploy is preserved via a branch-tip diff (before/after
for-each-ref) that fires cloud.OnGitPush for every advanced branch, and
metering (recordUsage) runs after every push. The go-git server transport
is removed; go-git remains only for bounded init/ref-read/object-building.
Runtime: add git to the alpine stage (upload-pack/receive-pack/http-backend/
git-remote-https).
Tests: real git-CLI clone+push round-trip, a 48 MiB clone proving the pack
streams with ~120 KB server heap growth, mirror from an external
git-http-backend source, and the existing tenant-isolation + push-to-deploy
suites, all green.
The v1.806.15 tag was re-pointed to the current fix commit; a stale local
module cache wrote the previous tag content's hash into go.sum in #290, so
cloud CI's fresh download failed verification (checksum mismatch / SECURITY
ERROR). Re-fetched the module clean so go.sum records the actual v1.806.15
content hash. No code change; go.mod pin unchanged.
Verified: CGO_ENABLED=1 go build -tags "libsqlite3 sqlite_fts5" ./cmd/cloud → 0.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Carries the hanzo-code-claude corrections (ai 68be82cf, first in v1.806.14):
per-upstream reasoning-effort vocabulary (GLM/DeepSeek max|high, not OpenAI
low|medium|high) and round-tripping assistant thinking as reasoning_content
so DeepSeek/Kimi tool-call loops no longer 400/stall. v1.806.15 also carries
ai's luxfi/geth proxy-resolution CI fix and the native Responses work.
Prod (cloud v1.799.16) embedded ai v1.806.13, which predates these fixes.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Add host-guarded root-level /:org/:repo/{info/refs,git-upload-pack,git-receive-pack}
reusing the existing handlers. onGitHost gates on the request Host == git host
(defaultSSHHost(Domain)); on api/console hosts the routes fall through (c.Next())
so a bare /:org/:repo never shadows another surface. Advertised clone URL stays
/v1/git until cutover. Test: TestRootSmartHTTP_HostGuard (git host → handler runs;
other host → 404).
zen mounts as a /v1-scoped Claim middleware BEFORE ai's /v1/* catch-all
(Wire position 100, ahead of ai at 150). Claim routes every request whose
model is a zen SKU to zen's serving layer in-process and c.Next()s the rest
to ai, so zen owns the zen family (identity, tools, 1M ladder, codec) and ai
owns every other model + /v1/models. The frozen wire sequence is updated.
Billing for zen* moves from ai's in-handler metering (skipped — zen claims
before ai's beego catch-all) to zen's own Gate+Meter wired to the same
commerce metering client the edge gate uses — ONE billing source for zen*,
never double-billed, never free. Granularity is org/project/user mirroring
the edge gate: home org (principal.BillingOrg) pays, project
(principal.ValidatedProject) scopes spend caps, user is attributed. An admin
acting in another org bills their own home org.
metering.Usage now carries a typed money.Amount (native 18-dp USD — the
finance ledger's own precision) as the canonical debit; Record passes it
straight to finance with NO flooring to cents or micros, so an exact
per-token cost from zen debits exactly. Legacy AmountCents/AmountMicros
reconstruct the same Amount for older callers. See hip-00NN.
Bumps hanzoai/zen v1.0.0 -> v1.1.0 (additive API: TenantResolver, ctx-aware
Meter/Gate, Tenant.BillingOrg/Project, Usage.RequestID).
zen5-ultra routes to anthropic-claude-opus-4.8, which 403s on this account.
Drop the ultra backend chain from the Fable tier so it resolves to zen5-pro
(the working heavy-reasoning tier) instead. Picks up ai v1.806.13:
upstream-429 surfacing (typed apiError, status reaches client) + the
Anthropic thinking-budget -> upstream reasoning-vocabulary adapter
(glm max|high, openai low|med|high) so deep CC thinking reaches the deep
tier and invalid vocab no longer gets GLM-5.2 429'd. See hip-00NN.
origin/main was rewritten (tracker/git work force-pushed over ~13 shipped
commits). This merge reunites both histories; go.mod resolved as the
highest-version union + go mod tidy.
The big.Int 18-decimal fixed-point implementation lives ONCE in the shared
packages; cloud/clients/money becomes the thin policy layer that pins the
credit unit to 18-decimal USD. Finance/treasury/admin call sites unchanged
in behavior — money, ledger, sqlstore, finance, admin-core tests green.
Correct the taxonomy after confirming the real architecture: helpdesk tickets,
CMS content, ERP docs and knowledge entries are framework.DocType records (Hanzo
Base), and CRM Company/Contact/Opportunity is a bespoke relational store — a
DIFFERENT plane from the tracker. The tracker Issue is the engineering/project
work-item primitive, not the universal record.
- kinds narrowed to issue|pr|epic (drop deal/ticket/doc = domain records; drop
task = the async plane, hanzoai/tasks). sources kept (legit work-item origins).
- contract.go rewritten to THREE orthogonal planes — tracker Issue (work items),
framework.DocType (domain records), hanzoai/tasks (async execution) — and the
thin one-directional seams (ExtRef links, Issue<->task), never embedding.
- Issue/IssueFilter/field docs + test seed aligned to the narrowed set.
Pin the alignment law as in-repo design source of truth (doc-file, like the
git package doc and tasks/CONTRACT.md): there is ONE work-item primitive — the
tracker Issue — and no second issue table. Every product surface (hanzo.team
board, git Issues/PRs tabs, CRM pipeline, helpdesk queue, CMS task list, an
agent's work) is a FILTER over the one table, never a parallel store.
Documents: the four immutable discriminators (Kind/Source/Repo/ExtRef) and
their orthogonality; the surface->filter mapping; the two tenancy roots (IAM
project = physical file; tracker Project = KEY-N team within it; git repo bound
by the Repo discriminator, a third thing); and the tasks seam — tracking
(intent+state) is not execution (durable async = hanzoai/tasks), composing in
exactly two directions (Issue->enqueue task, task->patch Issue board state),
never inverted.
Extend the tracker Issue with four immutable discriminators (Kind, Source,
Repo, ExtRef) so a project task, a git issue, a pull request, an epic, a CRM
deal and a helpdesk ticket are ALL the same row. Every product surface becomes
a FILTER over this one table — never a second tracker:
- hanzo.team board = ListIssues(Filter{}) (or {Status})
- git repo Issues = Filter{Repo, Kind:"issue"}
- git repo PRs = Filter{Repo, Kind:"pr"}
- CRM pipeline = Filter{Kind:"deal"}
- helpdesk queue = Filter{Kind:"ticket"}
store.go: additive kind/source/repo/ext_ref columns (DEFAULTed so existing
rows read as native team issues; idempotent ALTER for pre-spine DBs) +
org_repo / org_kind indexes; ListIssues takes IssueFilter (status+kind+repo+
source, each an optional bound predicate).
tracker.go: closed kinds/sources sets, normKind/normSource, issueFilter() from
?status=&kind=&repo=&source=; create accepts + validates the discriminators;
billing category stays the constant "issue" row, decoupled from work-item Kind.
Discriminators are identity — set once at Create, immutable on Update — so a
row never migrates between surfaces. Tests: TestIssuePolymorphicSpine proves
the filter slices + defaults + cross-repo source view.
The embedded git host gains its browser surface (ui.go + ui_templates.go),
server-rendered in the ONE cloud binary — the native replacement for the
standalone Gitea web app, so git.hanzo.ai can retire it. Routes at /git/*:
org repo list, repo home (branches/HEAD/tree/clone/README), tree browse, file
view (binary-aware), commit log. Reads the SAME org-scoped store + go-git object
storage the API uses; every page scoped to the IAM-validated X-Org-Id with the
path-vs-identity guard (no cross-tenant read). html/template auto-escaping is the
XSS boundary. Build+vet clean; all clients/git tests green.
The embedded Hanzo Git subsystem's env is now bare GIT_* (was CLOUD_GIT_*). Not
set in any prod deployment (code defaults), so the rename is behavior-neutral;
takes effect on next build. Cleaner + gitea-free naming.
The release has been failing on:
go: github.com/hanzoai/otel-collector@v0.144.10: unknown revision v0.144.10
That tag EXISTS on GitHub and resolves fine from a clean cache (verified:
go get github.com/hanzoai/otel-collector@v0.144.10 -> exit 0). The failure is
the BuildKit cache mount, not the pin: with no explicit id, BuildKit keys the
cache by target path alone, so a negative lookup recorded while the tag did not
yet exist is remembered FOREVER and there is no way to evict it.
Give every cache mount an explicit, bumpable id (cloud-gomod-v2 /
cloud-gobuild-v2). Bumping the suffix forces a cold cache. This unwedges the
release without touching the (correct) otel-collector pin, and gives us the
lever we were missing the next time a phantom pin poisons the cache.
Per directive: only OUR internal embedded git is rebranded (Hanzo Git), and the
default source provider is our own git.hanzo.ai ("git"). External providers
github/gitlab/bitbucket/gitea.com stay available as sources — gitea.com is fine
as an external source, we just never brand OUR host as Gitea.
The context window is a BODY-SIZE fact, not only a model fact: a chat request
carries its whole prompt in the request body. zip/fiber defaults BodyLimit to
4 MiB (zip.go: 'if cfg.BodyLimit == 0 { cfg.BodyLimit = 4 << 20 }') and cloud
never set it — so a 1M-token prompt (~4.3 MB of JSON) was refused by fasthttp
BEFORE any handler ran, with the opaque 400 'Error when parsing request' that
reads like a malformed payload rather than a size cap.
Effect: the 1M-context routes were unreachable in practice. Measured on prod
(v1.799.10, direct to pod:8000, bypassing ingress+gateway):
1.5 MB body -> 350,152 prompt tokens -> 200 (glm-5.2 correctly auto-rolled)
2.0 MB body -> 466,148 prompt tokens -> 200
4.0 MB+ -> 400 'Error when parsing request'
So the cascade logic was right all along; the transport silently capped it at
~900k tokens. 16 MiB gives a 1M-token prompt ~3.7x headroom. The edge is
authenticated + rate-limited and fasthttp streams, so this is not a memory
vector. Env GATEWAY_BODY_LIMIT, mirroring the GATEWAY_READ_BUFFER_SIZE
precedent (same class of bug one layer up: the 4 KiB header default).
Tests pin the invariant so it cannot silently regress to the framework default.
Closes the Studio→CMS→Commerce loop the other direction. The forward edge
(storefront.go) publishes a rendered Asset onto the product image; this adds the
REVERSE — a new catalog product triggers its ecom render — plus the join-key
integrity gate. Decomplected over the COMMERCE event stream so commerce and content
never import each other.
Reverse loop (B):
- events: SubjectProductCreated/Updated ("commerce.product.*", on the COMMERCE
stream) + PublishProductCreated/Updated on the existing events.Publisher (nil-safe
no-op when NATS is absent, mirroring the order publishers).
- commerce product REST: publishProductEvents middleware fires the catalog event after
a successful create/replace, reading the process Publisher off the gin context
exactly like the order publishers — a total no-op when unwired.
- clients/catalogsync: in-process consumer subsystem on the COMMERCE stream mapping
product.created → content.EnsureCatalogAsset (design == slug). Inert until
CLOUD_COMMERCE_NATS_URL names the NATS carrying the events.
- content.EnsureCatalogAsset: product slug → ONE ecom Asset draft, idempotent (skip if
a non-archived ecom asset already exists) and quiet (errNotConfigured / lane not
installed = skip, never an error loop).
Integrity (C):
- Storefront gains ProductExists (reusing the same commerce S2S seam, X-Org-Id-pinned);
enforceCatalogRefs before_save on Campaign.product fails closed on a dangling handle,
validates only on set/change, and skips when commerce is unconfigured or erroring.
Campaign-only on purpose: Asset.design is authored render-first (before the product
exists), so gating it would reject the loop's first step.
Wired after content in subsystems.Wire() (frozen order test updated). Tests cover the
EnsureCatalogAsset idempotency/skip paths, the Campaign dangling-handle rejection, the
ProductExists S2S wire, the catalog event envelope + no-op, the REST middleware
classification, and the consumer dispatch ACK/NAK semantics.
swap the 3 remaining driver call sites (audit_mirror, commerce/db/
datastore, o11y/event_ingest) hanzoai/datastore-go/v2 -> hanzo-ds/go,
and take o11y v1.5.26 (fully driver-clean). cloud graph now has zero
hanzoai/datastore-* — the driver's one home is hanzo-ds/go.
- ai v1.806.10: gofumpt-clean CI (unblocks ai release pipeline) +
deepseek-r1-distill context parity + swarm datastore standardization
(hanzo-ds/go v1.0.1, datastore-go v2.47.2).
- o11y v1.5.23 was a PHANTOM pin (tag never published; graph jumps
1.5.21 -> 1.5.25). Cloud only built via warm CI module cache; a fresh
resolve failed 'unknown revision v1.5.23'. Repin to the real latest
v1.5.25. Verified: go build ./cmd/cloud = exit 0.
The ONE interactive login: the CLI mints a device+user code from IAM
(hanzo.id /v1/iam/oauth/device, client hanzo-app — the first-party client
seeded with device_code), renders the verification link as text and a
terminal QR, and polls /v1/iam/oauth/token until approved. No password
ever touches the terminal; works headless (GPU boxes, ssh). Same endpoints
and client as hanzo-dev's live device flow, so one server-side config
serves both. --username/--password-stdin keep the password grant for
automation; --token unchanged.
detectGPUs only probed nvidia-smi, so a Metal/MPS box registered as
CPU-only in the fleet. On darwin/arm64 report the Apple chip as one GPU
with hw.memsize unified memory, MiB-formatted like nvidia-smi.
The gpu-jobs claim loop renders on the LOCAL studio server (127.0.0.1:8188);
naming a checkout makes connect own that server's lifecycle too — launch from
the venv, health-probe /system_stats, one grace re-check, free the port and
relaunch on death or hang. Replaces the GB10 watchdog script and hand-rolled
systemd units: the hanzo CLI is the one way a BYO box joins the fleet, render
backend included. --daemon bakes the flag into hanzo-gpu.service.
renames the store call site DatastoreDB() -> Datastore() (redundant DB
dropped) and bumps hanzoai/o11y v1.5.19 -> v1.5.23, which also completes
the clickhouse->datastore debrand indirects and pins a valid module zip
(v1.5.21/.22 zips were case-collision-invalid; .23 verified downloadable).
CGO_ENABLED=0 go build ./... green.
The double-entry ledger stored money as int64 cents, so every sub-cent AI
charge was floored to 0 (free AI) and every call skimmed the fractional
cent on rounding. Money is now EXACT to 18 decimals end to end.
- clients/money: new immutable big.Int Amount in atto-USD (1e-18) — the
EVM/ERC-20 uint256 unit, so the off-chain ledger and an on-chain credit
token are the SAME value with no boundary rounding. No float, no deps.
- ledger core + SQLite adapter: Posting/Entry/Balance int64 → money.Amount;
amounts stored as TEXT (atto overflows SQLite INTEGER past ~$9.20) with an
O(1) running-balance column and a one-time cents→atto ×1e16 rebuild
migration (existing prefunds carry over exactly on first open). Anchor
ComputeRoot hashes exact atto.
- finance wallet / types seam / Formance adapter / admin grant+deposit /
metering / treasury threaded through. Treasury reserve-fund facade stays
int64-cents (revenue-share is cents-granular).
- ai debit path (v1.806.6) emits exact decimal-USD; wireFinance parses it to
atto. Balance READ stays coarse cents (gates a >0 threshold only). Grant
response returns balanceAtto so a sub-cent debit is visible.
Pin ai v1.806.3 → v1.806.6 (exact nano-USD billing; datastore-go/v2 registers
driver 'datastore', no clash with cloud's hanzo-ds/go 'clickhouse').
A published content.Asset (kind in ecom/product/lifestyle, design==product slug)
materializes its S3 URL into the org's Hanzo Commerce store Listing headerImage —
the runtime display layer karma.style already reads (GET /v1/store/:store/listing).
Replaces the build-time studio->S3->library.json->sync-*->site pipeline with ONE
publish-edge side effect, decomplected exactly like the social Distributor.
- storefront.go: Storefront seam + commerce S2S impl (Bearer COMMERCE_SERVICE_TOKEN
+ X-Org-Id over clients/commerceinproc; store via GET /v1/store/current; upsert
PUT /v1/store/:id/listing/:design). Fail-closed (not_configured) with no token.
- content.go: wire sf edge at Mount; fire StorefrontPublish on the published edge;
TransitionResult.Storefront.
- Tenant-scoped by IAM org; assets referenced by S3 URL; no sync scripts.
Tests (CGO_ENABLED=0): gate, url resolution, transition side-effect, fail-closed,
and the real commerce S2S wire incl. X-Org-Id tenant pinning.
datastore-go was pinned at v2.47.1 — a version the module proxy cannot resolve (the /v2 module path does not exist at that tag), so cloud did not build from a clean cache. v2.47.0 resolves.
Claude-Session: https://claude.ai/code/session_01RFrWpXc1BsqfrFYMbyDusJ
billingSubject(org,name) = org (lowercased), always. Deletes the personalBillingOrgs
and orgBillingOrgs allowlist parsers (PERSONAL_BILLING_ORGS / ORG_BILLING_ORGS). Keeps
the console billing view in lockstep with ai/object.BillingSubject so the view and the
gateway gate scope to the SAME subject. Tests prove the killed envs are ignored.
Needs go.mod bump github.com/hanzoai/ai -> v1.806.8 (run in a module-resolving env).
datastore-go's transport moved to hanzo-ds/native, so the ClickHouse/ch-go
indirect is pruned from cloud entirely. No cloud package imports a ClickHouse
or signoz module.
- otel-collector v0.144.8 (+ replace => v0.144.8-hanzo.0) -> clean require v0.144.10
- dropped: replace SigNoz/signoz-otel-collector => hanzoai/signoz-otel-collector
and the // indirect SigNoz/signoz-otel-collector require
- cloud now pulls hanzo-ds/{go,native,mock} transitively; 0 signoz in go.mod/go.sum;
no cloud package imports a ClickHouse driver directly (build green)
Residual: one ClickHouse/ch-go // indirect graph line from a transitive dep's
go.mod (not compiled in) — clears when that dep drops it.
Measured DO context windows + config-only resolver (the old table refused
deepseek-v4-pro at 131072) + oversized-prompt reroute to the 1M model.
This is what lifts hanzo code claude to the full 1M context.
The o11y read plane was renamed signoz_* -> o11y_* (databases, tables, query
identifiers) at the source; the ClickHouse cutover migration
(o11y/deploy/clickhouse/migrations/0001_rename_signoz_to_o11y.sql) renames the
physical objects data-preservingly. Cloud's eval telemetry queried the OLD
physical name directly, so it would break post-cutover AND leaked the brand:
eval/metrics.go: spanTable signoz_traces.distributed_signoz_index_v3
-> o11y_traces.distributed_o11y_index_v3
Plus every prose/table-name reference across the in-repo o11y read plane and
commerce OTel bootstrap: signoz_traces/signoz_logs -> o11y_traces/o11y_logs,
SigNoz-fork provenance comments -> o11y/upstream. No source brand leaks remain.
DEPLOY: apply 0001_rename_signoz_to_o11y.sql during the collector cutover window
so prod ClickHouse objects match the new identifiers before this image serves.
Residual: go.mod keeps a REDIRECTED github.com/SigNoz/signoz-otel-collector key
(replace -> our fork; no SigNoz code fetched) — pulled transitively by
hanzoai/otel-collector's own go.mod; purge belongs to that fork's rename.
ai v1.806.6's UsageEvent replaced Cents(int64) with USD (exact decimal string,
atto-precise) so a sub-cent AI call bills precisely upstream. Adapt cloud's recorder:
- types.UsageInput gains USD (supersedes Cents when set).
- build.go SetUsageRecorder forwards u.USD.
- finance.RecordUsage rounds USD -> cents at the ledger boundary (usdToCents,
round-half-up via math/big, no float). The local finance ledger is cents-denominated,
so sub-cent floors to 0 exactly as the prior int64-Cents contract did — the atto path
is commerce, not this store. Money-safety locked by TestUsdToCents.
Also carries ai v1.806.6's glm-5.2 1M context window (already in v1.806.5). Build ./...
green, finance tests pass.
Claude Code treats `best` (like opus/sonnet/haiku) as a reserved model
alias and rewrites it to a claude-* id. api.hanzo.ai does not serve those
claude-* ids (a request 403s — see anthropicWire), so `hanzo code claude`
on the default `best` died at session start. Default to glm-5.2: the
stable GLM-5.2-class 1M-ctx frontier, a concrete catalog id CC passes
through unchanged; the backend still cascades on rate-limit / down.
The ledger already scoped correctly (org = which books, subject = which wallet in them), but both hooks passed the org for BOTH, collapsing every member onto the tenant's pool wallet. Since every signup lives in 'hanzo', a brand-new $0 account read HANZO's balance and sailed through the gate: we were enforcing our own wallet, not theirs. Now the gate reads, and usage debits, the subject ai already resolves — a person => their own wallet (personal plan), an org-owned application/service key => the org's account. That is the product: sign up as yourself with personal billing, then stand up an org whose users are your customers (Organization.Parent + AdministersOrg in hanzoai/iam). The invariant that must never break: the gate READ and the usage DEBIT key on the SAME wallet, or spend outruns the balance that admitted it — both use subject, keep them together. Also unpins o11y v1.5.16, whose tag was re-pointed upstream so its hash no longer matches go.sum (build fails verification); v1.5.17 is the unpoisoned tag.
Claude-Session: https://claude.ai/code/session_01RFrWpXc1BsqfrFYMbyDusJ
- tracker.go: prose 'Huly/Svelte' -> 'prior Svelte'
- model.json: Slack-mapping wire attrs hulyChannel/hulyChannelClass ->
teamChannel/teamChannelClass (seed-model data, NOT Go-referenced; valid JSON).
The clients/team/*.go debrand already landed on main. This closes it out.
NOTE: teamChannel wire IDs are new-workspace seed data; the deployed front
(front:v0.7.391) still speaks hulyChannel for Slack channel mapping, so a lockstep
front rebuild is needed for Slack mapping on NEW workspaces — flagged, not silent.
The ledger already scoped correctly (org = which books, subject = which wallet in them), but both hooks passed the org for BOTH, collapsing every member onto the tenant's pool wallet. Since every signup lives in 'hanzo', a brand-new $0 account read HANZO's balance and sailed through the gate: we were enforcing our own wallet, not theirs. Now the gate reads, and usage debits, the subject ai already resolves — a person => their own wallet (personal plan), an org-owned application/service key => the org's account. That is the product: sign up as yourself with personal billing, then stand up an org whose users are your customers (Organization.Parent + AdministersOrg in hanzoai/iam). The invariant that must never break: the gate READ and the usage DEBIT key on the SAME wallet, or spend outruns the balance that admitted it — both use subject, keep them together. Also unpins o11y v1.5.16, whose tag was re-pointed upstream so its hash no longer matches go.sum (build fails verification); v1.5.17 is the unpoisoned tag.
Claude-Session: https://claude.ai/code/session_01RFrWpXc1BsqfrFYMbyDusJ
Deployed glm-5.2 was capped at 16384 tokens (the context_length_util.go
fallback) — long Claude Code sessions 402'd. ai v1.806.5 sets glm-5.x to a
1M window with 131072 modern fallback (commit d248ff98).
The four Claude Code tier slots (Haiku/Sonnet/Opus/Fable) are a fixed
zen5-* capability contract, decoupled from the resolved main model id.
Previously OPUS tracked ANTHROPIC_MODEL, coupling the tier to the main
choice; now OPUS=zen5-pro and FABLE=zen5-ultra are stable.
Decomplects which-tier from which-model: the tier->alias map is fixed
in the client; the alias->upstream map lives in models.yaml. Swap an
upstream (GLM -> Qwen 3.6 -> a future frontier) and every SDK/CLI/CC
integration keeps working unchanged. zen5-ultra carries its own backend
fallback chain (zen5-ultra -> zen5-pro -> zen5 -> zen5-flash) so the
Fable tier degrades gracefully if the premium upstream is unavailable.
claude code's subagents, permission classifier, and /compact default to
built-in claude-* model ids (claude-haiku-*, claude-opus-*, claude-sonnet-*).
api.hanzo.ai does not serve those ids — a request 403s, which:
- kills the classifier ('auto mode cannot determine safety of Bash')
- kills every subagent (the 403 that dead-ended session cff690fc)
- leaves /compact to run on a non-1M model -> 262145 > 262144 -> unresumable
anthropicWire now pins every CC tier slot to a served zen5 alias (the
Hanzo-standard mapping of CC tiers onto top OSS models resold via DO GenAI):
ANTHROPIC_MODEL = <model> (default best -> zen5/glm-5.2)
ANTHROPIC_SMALL_FAST_MODEL = zen5-flash (DeepSeek-4 Flash, classifier)
ANTHROPIC_DEFAULT_HAIKU = zen5-flash
ANTHROPIC_DEFAULT_SONNET = zen5 (GLM-5.2)
ANTHROPIC_DEFAULT_OPUS = <model>
ANTHROPIC_DEFAULT_FABLE = zen5-pro (DeepSeek-V4 Pro)
Live-proven: zen5/zen5-flash/zen5-pro/zen5-ultra all return 200 on the account;
only the raw claude-* ids 403. tests: TestAnthropicWirePinsZen5Tiers +
TestAnthropicWireExplicitModel lock the mapping and forbid raw claude-*.
Debrand the team subsystem (our fork is Hanzo Team, not Huly):
- prose/comments: Huly -> Team / the platform
- hulyName() -> personName() (internal Person.name formatter)
- HULY_MODEL_VERSION -> MODEL_VERSION (no brand prefix at all, per the naming rule)
- TestHulyName -> TestPersonName
model.json WIRE class IDs (slack:class:SlackChannelMapping_hulyChannel, attrs
hulyChannel/hulyChannelClass) are deliberately UNTOUCHED: the deployed front
(ghcr.io/hanzoai/front:v0.7.391) speaks them, so renaming needs a lockstep front
rebuild — tracked separately, not a silent prod break.
Build + team tests green.
Bump hanzoai/ai to the ORG_BILLING_ORGS build (v1.806.3 + the allowlist), so the
gateway's BillingSubject promotes an allowlisted org (e.g. hanzo) to ONE shared
pool. Mirror the same scoped override in the console billing BFF (billingSubject)
so the console view scopes to the SAME subject the gate reads. Default empty =
zero behavior change.
Also refresh the stale hanzoai/o11y v1.5.16 go.mod checksum in go.sum (metadata
only; the zip/code hash was already correct — the readonly build passes), which
a moved tag had left inconsistent and which blocked every cloud build.
Change defaultCodeModel glm-5.2 -> best so `hanzo code claude` (no model arg)
auto-routes to the best-available coding model by quality and cascades on
rate-limit / out-of-credit / down (server-side, controllers/failover.go).
Explicit overrides still work (`hanzo code claude glm5.2`); the fuzzy resolver
validates `best` against /v1/models, where it is a real listed catalog entry.
Completes the datastore standardization for cloud's own source: audit_mirror,
commerce/db, o11y/event_ingest now import github.com/hanzoai/datastore-go/v2
(the rebranded fork registering the UNIQUE sql driver name "datastore") instead
of the interim github.com/hanzo-ds/go (which registers "clickhouse" and collides
with upstream). Bumps ai v1.806.3 → v1.806.4 (its object/datastore.go likewise on
datastore-go/v2). cmd/cloud now links ZERO upstream ClickHouse/clickhouse-go/v2.
Remaining hanzo-ds/go is transitive via o11y/pkg/datastoremetrics — needs an o11y
release built off datastore-go/v2 (v1.5.16 tag currently resolves to the interim
hanzo-ds/go variant).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
An in-process metering/billing dispatch (cloud → commerce via buildMeteringClient)
carries the verified COMMERCE_SERVICE_TOKEN *and* an X-Org-Id header (to select the
tenant namespace). The X-Org-Id makes IAMTokenRequired stamp iam_authenticated=true
with ZERO permissions (an S2S call carries no X-User-Permissions), so on an
Admin-masked billing route (/v1/billing/{balance,tier,usage}) TokenRequired's IAM
branch ran first, hasScope(0, Admin)=false → 403 "IAM principal lacks required
permission scope". The service token was never consulted; the balance read / usage
debit failed and the completions edge rendered that internal failure as a client 500.
Fix: in TokenRequired, check the service-token branch BEFORE the IAM branch. A request
bearing the KMS-sourced service token is the trusted platform S2S caller and must
authenticate as the service, never be subjected to the per-user IAM scope gate — even
when it also carries X-Org-Id. The two branches' logic is byte-identical; only the
order changed.
- Does NOT loosen the gate for external callers: a real IAM user never holds the
service token, so they still hit the IAM branch and are masked-gated exactly as
before (TestTokenRequired_UserIAMScopeGateUnchanged, IAMBranchEnforcesMasks).
- Prepaid gate stays fail-closed: funded → allowed, unfunded → clean 402
insufficient_balance (TestInProcMeteringDispatch_ServiceTokenAuthPath drives the
real chain over commerceinproc with finance NOT co-resident).
- The service token is KMS/env-sourced and never logged.
The 500-render itself lives in the hanzoai/ai dep (it serves /v1/chat/completions and
turns the internal balance/debit failure into a 500); once the cloud-side dispatch no
longer 403s, that failure no longer occurs.
The ai gate read the account pool but debited the per-user wallet (read-subject != debit-subject),
so a funded account gated but never depleted. Key BOTH the balance read and the usage debit on the org
(its default-account pool wallet) in the wireFinance hook, so a funded account gates AND meters
consistently. Fund accounts, not users.
* feat(cron): ONE durable platform cron on the tasks engine — retire every ticker and k8s CronJob
clients/cron replaces the k8s CronJob fleet AND the commerce sweep ticker
with durable schedules on the shared embedded hanzoai/tasks engine
(v1.50.0 — the release whose sweeper actually fires on the sharded store).
The engine owns time; runs are durable workflows visible in the Tasks
console (/_/tasks, tasks.hanzo.ai) under the CRON_ORG shard (default
hanzo), namespace default, queue cloud-cron.
Entries are DATA in universe git: a ConfigMap labeled
cron.hanzo.ai/enabled="true" carries schedule + EITHER job.yaml (a
batchv1 Job manifest run to completion with Forbid concurrency, entry
label, TTL self-reap) OR poke.json ({url,method,bearerEnv,timeout} —
bearer resolved from THIS process's KMS-synced env, never stored).
A reconcile workflow — itself a durable schedule (*/5) plus one boot
pass — diffs ConfigMaps against cron-prefixed schedules: upserts are
drift-gated (rewriting resets the fire anchor), deletes never touch
foreign ids. Fire-time activities re-read their ConfigMap, so payload
edits apply next tick.
The commerce auto-recharge sweeper (clients/commerce/sweep.go) is
deleted; its 15m POST /v1/billing/auto-recharge/run-all becomes the
cron-billing-autorecharge poke entry — same request, same token, one
cron system.
Tests (-race): TestPokeEndToEnd + TestJobEndToEnd drive the FULL durable
path against a real embedded engine (org-shard schedule → trigger →
loopback worker → activity → completed run in the org shard — the exact
console-visibility + dispatch-routing the design stands on);
TestReconcileConverges pins upsert/delete/foreign-id/anchor semantics;
TestParseEntry pins the ConfigMap contract. commerce + subsystems suites
green.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
* feat(tasks): UI at /tasks (no /_/) + org/project/user-native gate (tasks v1.51.0)
The Tasks console moves to console.hanzo.ai/tasks + tasks.hanzo.ai/tasks —
the /_/ prefix is gone. Subsystem routes mount before the console SPA
catch-all, so /tasks wins the path and every other console route keeps the
SPA fallback. The gate now threads the FULL validated identity — org,
project (X-Project-Id, minted by identity middleware from the validated
project claim), user, email — into the engine via tasks v1.51.0's
WithIdentity; convention: project ↔ tasks namespace inside the org shard.
ZAP plane unchanged (loopback ungated + IAM-gated cluster listener).
The embedded SPA bundle is rebuilt with base /tasks/ in a follow-up sync
commit (hanzoai/admin builds it; clients/tasks/ui/dist is the embed).
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
* ci: retrigger (Actions dropped the push event)
* chore(tasks/ui): sync SPA rebuilt with base /tasks/ (hanzoai/admin@448b189)
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
---------
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Prefer the datastore-named env keys, fall back to the clickhouse-named legacy
keys so a rollout mid-flight (new binary / old manifest, or vice versa) never
drops the audit OLAP mirror connection. Code-only; no roll.
ai v1.805.10 pulled upstream ClickHouse/clickhouse-go via ai/object, which double-registered
the clickhouse driver against cloud's hanzo-ds/go fork → boot panic (CI smoke caught it). ai v1.805.11
merges the clickhouse→hanzo-ds purge, so the upstream driver is no longer imported (verified: 0 in the
build, binary boots clean). Also: commerce.BalanceCents (native in-proc balance read — backfill + cockpit
read real balances, no commerce.inproc) and POST /v1/admin/finance/deposit (fund a SPECIFIC wallet:
hanzo/z, not just the pool).
LOW-1 fast-follow (Red). In these two create-handlers the DEBIT correctly keyed on the
HOME org (principal.Payer(c)) but the pre-create balance GATE still keyed on the
EFFECTIVE org — the two swapped only the meter, not the gate (their Gate signature had
gained a projectValidated arg, so the class swap missed them). Effect: a SuperAdmin
masquerading into a victim org was balance-gated on the VICTIM's funds while the debit
landed on admin's ledger (money already correct; an availability/consistency nit).
Fix: swap the Gate org arg -> principal.Payer(c) in both handlers, restoring the
invariant the other ~8 resource meters hold — gate + debit BOTH key on home; data scope
(namespace/project) stays effective.
Test: TestResourceMeter_GateKeysOnPayerForMasqueradingAdmin resolves principal.Payer(c)
from a masquerade ctx (X-User-Owner=admin, X-Org-Id=victim -> admin) and asserts the
recCommerce balance check keyed on admin (home), never victim. go build ./...=0 (links),
gofmt clean, root ResourceMeter suite + clients/ml green (provisioning store tests are the
pre-existing CLOUD_KMS_MASTER_KEY_REF env gate, orthogonal).
Ports the prior Git-over-SSH + ZAP transport + client-less REST push work
(commit 95ba376, written against the old *svc-receiver service) onto the
current cloud.Service[state] generics framework on main. A port, not a
rewrite: the DRY architecture is preserved — ONE control-plane core, ONE
git pack path, three thin transports (REST, ZAP, SSH) over them.
Handler pattern adapted: every `*svc` method became a free function taking
`s *cloud.Service[state]` (matching git.go's existing create/list/get/del),
`cloud.TenantStore` → `cloud.NewOrgStore`, methods → `cloud.Handle(s, fn)`
route registration, `s.log`→`s.Log`, `s.stores/storage/keys/ssh`→`s.State.*`,
`tenant(c)`→`org(c)` (principal.Org). Package-scoped helpers (storeFor,
provision, recordUsage, refState, cloneURL, sshURL, firePushBuilds, session)
are the current framework's free-function forms.
New files:
- core.go coreCreate/List/Get/Delete/Usage — the ONE transport-agnostic
control-plane impl; REST + ZAP both call it.
- pack.go serve{Upload,Receive}Pack + ssh{Upload,Receive}Pack — the ONE git
pack code path both smart-HTTP and SSH drive (io.Reader/Writer,
transport-agnostic); readerOnly guards the SSH channel write side.
- ssh.go golang.org/x/crypto/ssh server. Host key from CLOUD_GIT_SSH_HOST_KEY
(KMS env) or on-disk ed25519 generated+persisted 0600.
PublicKeyCallback resolves key→(org,user) by SHA256 fingerprint,
fail CLOSED. Session accepts only git-upload-pack/git-receive-pack,
enforces path-org == key-org. Listen CLOUD_GIT_SSH_ADDR (:2222).
- keystore.go fingerprint-indexed global SSH public-key registry (public keys
only — nothing to hash).
- keys.go POST/GET/DELETE /v1/git/keys.
- push.go POST /v1/git/repos/:name/push — build tree+commit from posted
utf-8/base64 files, advance ref, create-on-first-push (composes the
ONE provision), fire the SAME firePushBuilds hook receive-pack fires,
return commit + cloneUrl + sshUrl.
- zap.go git/zap/{createRepo,listRepos,getRepo,deleteRepo,usage} envelope
adapters over the SAME core funcs, on the shared /zap plane. No
second ZAP server, no gRPC.
Modified:
- git.go repoView/toView gain sshUrl; state holds sshHost/ssh/keys; Mount
opens the keystore + starts the SSH listener; routes() wires
push+keys+ZAP; Shutdown stops SSH + closes the keystore. REST
control-plane handlers are now thin adapters over core.go.
- smart_http.go uploadPack/receivePack drive the shared pack.go path via
resolvePackRepo; firePushBuilds + session helpers retained.
Tests (go test ./clients/git/... — 16 pass): SSH key register accept/reject +
cross-org reject + end-to-end SSH clone+push over an in-process listener; a
zapface WS round-trip proving the ZAP path hits the SAME core as REST; REST
push create/update/build-hook, first-push provisioning, validation, and
cross-org isolation. tenant_isolation_test + git_test bind the SSH listener to
an ephemeral loopback port so tests never collide on :2222.
CGO_ENABLED=0 GOWORK=off go test ./clients/git/... passes; go build clean;
gofmt clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
metering.Client fetchAvailable/Record now resolve finance.Current() first (micros→cents
ceil so metered_ai's AmountMicros debit is never dropped), commerce HTTP only when finance is
not co-resident. admin ApplyGrant credits the finance wallet (fixes the broken commerce.inproc
grant path). New DepositInput.Ref makes deposits idempotent; finance.MigrateOrg + SuperAdmin
POST /v1/admin/finance/backfill?org= carry each org's commerce balance into its finance wallet
exactly once at cutover. All money — balance, usage, deposit — now flows through the ONE per-org
double-entry ledger.
Verified: build + go vet clean; finance/metering/admin/core tests pass (micros fold, ref
idempotency, backfill exactly-once, zero-HTTP-when-co-resident all locked by tests).
New clients/finance: ONE lightweight SQLite file per customer (orgs/<org>/finance.db),
double-entry wallet on the treasury/ledger core — deposit=funding→wallet, usage=wallet→revenue,
balanced within the org file, idempotent. Orthogonal money interface types.FinanceClient
(finance.Client, mirrors commerce.Client); NOT bolted onto the Deps bag and NO pick/RPC/disabled
cruft — a narrow published seam (finance.Publish/Current) every money consumer resolves.
BuildDeps.wireFinance constructs+publishes it and installs the ai router's typed
BalanceReader/UsageRecorder hooks (ai v1.805.10) so the prepaid gate reads balance + debits usage
by a DIRECT in-proc call, no HTTP. Rob-Pike small interfaces between systems, not a god-struct.
Remaining before cutover: edge meter + admin grant onto finance.Current(), backfill commerce
balances into <org>:wallet, then build/deploy/prefund/drop BALANCE_EXEMPT_USERS.
Root cause of the jammed lane (failed releases + phantom tags like v1.786.221):
the push step published the racy v<X.Y.Z> tag BEFORE it was confirmed free, and
the tag step hard-failed when a concurrent release had claimed that number
(computed 220, already tagged → 'rejected, already exists'). Result: a pushed
image with no matching git tag, mutable :vX clobbering, and dead releases.
Fix — decouple the immutable artifact from the versioned release:
- Push step publishes ONLY the unique sha-<sha7> (+ floating latest); never the
racy version tag.
- Tag step assigns the next FREE v<X.Y.Z> ATOMICALLY: recompute fresh, retag the
proven sha-image → :vX/:X.Y.Z/:X.Y via imagetools (metadata-only, byte-identical),
git-tag it, and RETRY on collision (the git-tag push is the serialization point —
a loser recomputes). Concurrent releases each grab a distinct free number.
- Compute-step collision is now a non-fatal hint (Tag step owns assignment).
- notify-universe rolls the Tag step's FINAL version, not the compute guess.
Preserves the invariant: git tag vX ⇔ image :vX pushed + smoke-passed. Logic
unit-tested locally (free-find + collision-skip). YAML valid.
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
THE billing fix (B): split 'who pays' (HOME org, X-User-Owner) from 'whose data'
(EFFECTIVE org, X-Org-Id), which the old code conflated onto one org. A platform
SuperAdmin masquerading into another org now spends from the admin ledger; data
scope stays on the acted-on org.
- clients/principal: Owner(c) (X-User-Owner, bounded+cloned), BillingOrg(c) (home w/
effective fallback, Validated-gated), Payer(c) (bare-string for the in-handler meters).
- middleware_identity: mint X-User-Owner=home for every validated principal (before the
admin org-switch) + strip it on ingress (authorityHeaders) so a client can't forge who pays.
- middleware_billing: identityFromCtx keys User+Org (balance check AND debit) on BillingOrg.
- 11 resource-meter live sites (visor/s3/security/platform-run/functions/ml/provisioning/
bots/tracker/automations): Gate+Meter+MeterUsage bill principal.Payer(c) (home); Org stays
effective for the namespace. Background/reconcile paths (a.Org/r.Org/b.Org, meterRun) and
the studio_render 'org' param bill the RESOURCE's own org — FLAGGED (see Red handoff).
- metered_ai + types: ChatRequest/EmbedRequest gain BillingOrg; meteredAI gate+record key on
billedOrg(BillingOrg,Org); threaded principal.Payer(c) through the code engine (Synthesize/
Embed) so internal RAG bills home too. Inner AI call keeps req.Org (data scope).
Part A (embedded commerce, clients/commerce mirror of the commerce repo): drop the spoofable
IsSuperAdmin boolean; auth.IAMClaims.IsSuperAdmin() gates homeOrg()=="admin" (HomeOrg from
X-User-Owner ?: Owner); iammiddleware reads X-User-Owner->HomeOrg (Owner stays effective);
EdgeAuth strips+mints X-User-Owner before the ?org override, stops minting X-User-IsSuperAdmin;
all .SuperAdmin() call sites -> .IsSuperAdmin(). Org-scoped IsAdmin untouched.
Tests (green): TestIdentityFromCtx_AdminMasqueradeBillsHomeOrg (owner=admin+effective=victim ->
debit on admin, data on victim), TestBilledOrg, TestMeteredAI_AdminMasqueradeBillsHomeOrg
(end-to-end debit user=admin), vendored TestIAMClaims_IsSuperAdmin masquerade+anti-escalation,
EdgeAuth strips forged X-User-Owner. go build ./...=0, go vet=0.
Rollout: pre-gateway (no X-User-Owner) billing falls back to effective (home==effective for a
normal caller; masquerade fails CLOSED). Deploy gateway (mints X-User-Owner) first.
Three defects found by testing a real completion instead of --version:
- an unreachable catalog silently passed the typed id straight through, so an
agent booted on a model that does not exist and failed opaquely mid-session;
resolution is now strict.
- a stale ANTHROPIC_API_KEY in the shell outranks ANTHROPIC_AUTH_TOKEN and
silently wins; it is cleared for the claude wire.
- codex ignores OPENAI_BASE_URL and talks to chatgpt.com unless a provider is
declared; declare Hanzo and select it (wire_api=responses, per upstream).
Cloud was the odd one out: it minted SuperAdmin from TWO signals
(claims.IsAdmin && owner == adminOrg) while IAM's canonical User.IsSuperAdmin() is
just user.Owner == conf.AdminOrg, and IsOrgAdmin folded SuperAdmin into itself. Two
predicates that each meant one-and-a-half things.
Decomplected — two orthogonal facts, one predicate each:
IsSuperAdmin = owner == adminOrg (platform sudo; the SAME equality IAM uses)
IsOrgAdmin = the IAM isAdmin bit (admin of one's OWN org; implies nothing about super)
A gate admitting either now writes IsSuperAdmin(c) || IsOrgAdmin(c) explicitly, so the
superset is visible AT the gate, not hidden inside a predicate. GuardScoped already did.
The admin org holds ONLY SuperAdmins (provisioned in, never promoted), so membership IS
the fact — the isAdmin bit is the orthogonal org scope, never a super gate. The KMS
machine-principal exclusion STAYS (a real guard: an admin-org machine token must never
be super). Test renamed + strengthened to lock the one predicate. Build ./... clean,
identity/admin/principal suites all green, KMS-machine exclusion tests pass.
Advances the in-process IAM embed (clients/iam, HIP-0106) from the
v1.31.20 pre-release pin (31a798fb, #117) to the published v1.31.22 tag
(387c4a60). Brings the project claim mint + default-project seed (#120)
into cloud's embedded IAM so console admin surfaces match standalone
hanzo.id, already on v1.31.22.
API-compatible: iamserver.InitEmbed and object.Project CRUD unchanged
(iam go.mod hash identical across versions), no clients/iam adaptation.
CGO_ENABLED=0 go build ./... green; clients/iam + clients/platform tests pass.
zip v1.6.0 deleted the deprecated/redundant surface (App.Mount/Route/ModuleFn/
UseFiber). cloud is off App.Mount (#257); licensing v0.1.4 migrated its one
App.Mount call. Build 0, framework tests green.
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
One way to start any coding agent on any Hanzo model: endpoint, credential and
model are injected, so nothing is exported by hand. Two wires cover all three
agents (claude speaks Anthropic; codex and @hanzo/dev, a codex fork, speak
OpenAI), model ids resolve fuzzily (glm5.2 -> glm-5.2), and agents run full-auto
unless --safe. Config gains apiKey; the key the rest of the toolchain already
keeps in ~/.hanzo/config.json is read (never written) so one key serves every tool.
WS4 audit flagged projects.Project and tracker.Project as duplicate types.
They are not: projects.Project is the Slug-keyed deployable site (one global
projects.db, org column); tracker.Project is a KEY-prefixed Linear-style issue
team living in a per-(org,project) tracker.db WITHIN one of those sites. Only
the universal storage-row skeleton (id/org/name/description/timestamps) is
shared; the natural keys carry different semantics (DNS label vs issue-ID
prefix) and the field sets are not subsets. Record the distinction on each type
so the finding is not re-litigated. Comment-only; no API or behavior change.
Co-authored-by: Hanzo <dev@hanzo.ai>
principal.ValidatedProject(c) now flows through ResourceMeter.Gate into
metering.AuthInput.ProjectValidated, so resource-creation caps harden
consistently with the edge BillingGate. Removes the hardcoded
ProjectValidated:false at resource_billing.go.
- Gate signature gains projectValidated bool (adjacent to project, mirroring
principal.ValidatedProject's return + AuthInput's field order).
- Every *zip.Ctx caller passes principal.ValidatedProject(c); the three
no-principal/client-body paths (metered_ai LLM decorator, background agent
run, content studio in.Project) pass false — unvalidated stays soft, never
fabricated true.
- Meter/Record path unchanged: Usage carries no ProjectValidated; cap
enforcement is Gate-only, so threading it there would be dead code.
Stays SOFT in prod today: ValidatedProject is true only for a validated NAMED
project, and IAM seeds none yet, so no org has a project claim. No behavior
change now; named-project caps auto-harden per-org as IAM seeds them.
Co-authored-by: Hanzo <dev@hanzo.ai>
The build path (buildFrontendCmd) already chose its BuildKit frontend on
dockerfile presence alone: an explicit Dockerfile → dockerfile.v0, otherwise
hanzoai/pack via gateway.v0. buildType never gated the build — it was stored
metadata whose closed set still advertised the retired nixpacks/buildpacks/
static strategies and defaulted git apps to nixpacks.
Collapse it to the one true surface:
- closed set is {pack, dockerfile}; git source defaults to pack, not nixpacks
- image source always yields buildType image (no build), forced by source, so a
client can no longer stamp a git strategy on a prebuilt-image app
- Goa contract (design.go + generated openapi3.json/yaml) enum, description and
examples track the same {pack, dockerfile, image} set
Adds TestBuildTypeSurface pinning the default (pack), the dockerfile escape
hatch, rejection of every retired strategy, and image-source forcing.
Co-authored-by: Hanzo <dev@hanzo.ai>
Six identity reads re-derived org/project inline instead of going through
principal — the ONE accessor. Route them all through it so the trust
decision lives in exactly one place and cannot drift.
- analytics tenant(): delegate to principal.Org — fixes a retained-buffer
bug (org keyed the cloud_usage ledger past request end as a zero-copy
fasthttp view; principal.Org clones). Gate unchanged.
- git/security projectScope(): read via principal.Project; absent header
and literal "default" both map to the un-suffixed default scope via
principal.IsDefaultProject, keeping today's (org,project) key shape.
- admin core.ResolveScope / core.GuardScoped / me / core.EmitAudit: read
org via principal.Org (validated principal + org), composed with the new
principal.IsOrgAdmin predicate. Admin-org bucket kept local.
Co-authored-by: Hanzo <dev@hanzo.ai>
Pin hanzoai/ai v1.805.7 (exempt concept removed: no BALANCE_EXEMPT_USERS/KEYS, no
fail-open; balance-unverifiable now DENIES). BuildDeps sets aiobject.CommerceTransport
= commerceinproc.Transport() and self-configures the ai module's commerceEndpoint to the
in-process placeholder when commerce is co-resident, so the embedded ai router's balance
read + usage debit hit the raw commerce gin (service-token, no socket) — the SAME one
ledger the metering client bills — not the customer proxy that 401s a service token.
Every external inference is now prepaid-gated with zero exempt principals.
hanzo/z was a normal customer, never admin; exempting it leaked money. Closed.
Projects are now owned by Hanzo IAM (hanzo.id) as the org-scoped
(owner,name) resource. Platform REFERENCES that store instead of owning
one: apps still live under a project, but create/list/get/delete/exists
of the bare project delegate to IAM in-process.
- New ProjectStore port (projects.go) with an in-process iamProjects
adapter over github.com/hanzoai/iam/object — no HTTP hop, and IAM's
canonical *object.Project is used verbatim (no platform-local clone).
- Delete platform's project ownership: the Project struct, the
platform_projects table, and Create/Get/GetByID/List/Update/Delete
project store methods. DeleteProject is replaced by DeleteProjectApps
(cascade of the app tree only; the project row is IAM's).
- Apps key on the IAM project NAME (platform_apps.project_id holds it);
the internal proj_ id indirection is gone, so reconcile/preview/domains
use app.ProjectID directly and drop their project lookups.
- run.go get-or-creates the default project via the ProjectStore, using
principal.DefaultProject as the single source of truth for "default".
- tenant() gates on principal.Validated(c) (canonical), not a raw
c.User()=="" whitespace-blind check.
- Tests: fake ProjectStore + TestProjectLifecycleDelegatesToIAM proving
create/delete route through IAM and delete cascades platform's apps.
Co-authored-by: Hanzo <dev@hanzo.ai>
SanitizeIdentity now mints X-Project-Id from the validated JWT `project` claim
(idClaims.mintedProject), exactly like X-Org-Id from `owner` — the raw client
X-Project-Id is never a source. Still checked non-foreign to the effective org, so
a global admin viewing another org drops their own-org project. Mirrors the edge
(iamauth.Claims.MintedProject), so the in-binary path binds the same header.
principal.ValidatedProject flips CONDITIONALLY: (project, true) only for a
validated, NON-default project claim — the signal a project-scoped spend cap uses
to HARD-enforce. The default project stays soft (IAM does not seed default
projects yet; a blanket true would wrongly hard-402 every org). Named project caps
auto-harden one org at a time as projects are seeded.
Tests: claim-sourced project binding + forged/override client X-Project-Id ignored
+ cross-org claim refused on admin org-switch + X-Org-Id forgery still blocked;
ValidatedProject claim-backed→(project,true), absent/default→soft, unvalidated→soft.
Follow-up (WS3, out of scope): resource_billing.go Gate hardcodes
ProjectValidated:false — resource-creation caps stay soft until it threads
principal.ValidatedProject.
Co-authored-by: Hanzo <dev@hanzo.ai>
Eliminate "global admin" terminology across the cloud repo in favor of the
standard SuperAdmin term. Pure naming refactor; behavior preserved.
Identifiers (commerce auth + middleware + root):
- (*auth.IAMClaims).GlobalAdmin() -> SuperAdmin() + all call sites
- IAMClaims.IsGlobalAdmin field -> IsSuperAdmin (JSON tag isGlobalAdmin
-> isSuperAdmin; SuperAdmin() still honors owner=="admin", the canonical
predicate, so behavior is preserved)
- edgeauth isGlobalAdmin() helper -> isSuperAdmin(); requireGlobalAdmin ->
requireSuperAdmin; test-local globalAdmin vars -> superAdmin
- const HeaderUserIsGlobalAdmin -> HeaderUserIsSuperAdmin, value
"X-User-IsGlobalAdmin" -> "X-User-IsSuperAdmin" (commerce-internal header:
minted+stripped in edgeauth, read in iammiddleware; no external sender)
- auth/globaladmin_test.go -> auth/superadmin_test.go
clients/admin: the deprecated isGlobalAdmin JSON alias (sibling of the
canonical isSuperAdmin, same value) is REMOVED rather than renamed — a rename
would collide with the existing IsSuperAdmin field, and the alias was a
transitional back-compat scaffold for this very migration. Tautological
alias-equality assertions dropped; isSuperAdmin checks retained.
Prose: "global admin" / "global-admin" / "GLOBAL-ADMIN" / "GlobalAdmin" in
comments and docs (LLM.md, k8s manifests) rewritten to SuperAdmin, case-aware.
Left untouched: cloud gateway header X-User-IsAdmin; principal.IsSuperAdmin /
IsOrgAdmin; svcorg's DefaultNamespace (global) kind.
build ./... clean; vet clean; SuperAdmin behavior-locking tests pass
(commerce auth/edgeauth/iammiddleware/catalog/costs/checkout, admin
scope/money, principal, root identity). gofmt clean. Zero GlobalAdmin /
global-admin remaining in .go/.md/.yaml.
Two SuperAdmin endpoints for multi-provider credit-management (the console renders
them; contract is authoritative):
GET /v1/admin/providers/credit -> [{provider, grant_cents, burn_cents,
remaining_cents, runway_days, has_credit, is_paid_only}]
GET /v1/admin/usage/funding?from&to -> [{provider, model, funding, tokens,
cost_cents, requests}], funding in {credit,paid,paid_only,byo}
DRY: reuses the admin auth guard (core.Guard), finance.go's DO billing read (real
k grant, live remaining/burn/runway), and the ONE cloud_usage warehouse (no new
store). DO is the seeded real row; others land as keys arrive. Funding is
provider-level-derived in v1; the per-call split lands when the ai metering write
stamps a funding column (next commit). Envelope = core.OK {status,msg,data:[...]}.
Two-ledger model kept distinct: UPSTREAM provider credits here vs DOWNSTREAM Hanzo
customer billing (commerce). Pure logic TDD'd.
Per the SuperAdmin convention (SOC2/FedRAMP; standard term SuperAdmin, NEVER
'global admin'): name the two admin scopes explicitly.
- principal.IsSuperAdmin(c) = c.IsAdmin() — platform sudo; SuperAdmin ⟺ owner==admin org
(X-User-IsAdmin minted only for that identity).
- principal.IsOrgAdmin(c) = IsSuperAdmin(c) || X-User-IsOrgAdmin — admin of one's own org.
GuardScoped now reads IsSuperAdmin (fast-path) + IsOrgAdmin (scoped) instead of bare
c.IsAdmin()/OrgAdmin. Pure rename — no behavior change, identity suite green.
v1.805.6 makes credit/quota exhaustion + agreement gates (402/403/insufficient_quota)
cascade to the next provider in the fallback chain — Phase 2 of multi-provider
credit-management + fixes the do-ai->anthropic opus fallback.
GuardScoped admitted ANY validated org member to their own org's admin
panels (/v1/admin/{overview,orgs,users,usage,analytics,me,bases}), not
just org-admins: SanitizeIdentity discarded the JWT org-level isAdmin bit
for non-admin-org principals (it only minted the GLOBAL X-User-IsAdmin
when owner==adminOrg), so GuardScoped's fallback (validated User+Org) let
a plain member through.
Mint a new, unforgeable X-User-IsOrgAdmin signal and require it:
- middleware_identity.go: add X-User-IsOrgAdmin to authorityHeaders
(stripped on every ingress request, so a client can never forge it),
and mint it for any validated principal with claims.IsAdmin &&
!isKMSMachinePrincipal — covering both a global admin and an org-admin,
owner-scoped and machine-excluded.
- clients/principal: add OrgAdmin(c) = IsAdmin() OR X-User-IsOrgAdmin.
- clients/admin/core.GuardScoped: require principal.OrgAdmin(c) in
addition to validated User+Org; a validated non-admin member now gets
the same 403 as an unvalidated caller. Guard (global-only) untouched.
Tests: orgAdminHdr now carries the bit; add TestScope_MemberWithoutOrg
AdminDenied (member without the bit is 403 on every scoped panel) and
TestSanitizeIdentity_OrgAdminHeader (org-admin gets the bit not global,
member/machine get neither, forged bit stripped on bearer + anon paths).
Finding 3 (grant durability — ARCHITECT: FAIL-CLOSED). EmitAudit is a silent
no-op when State.AuditStore == nil, so on a no-audit-store deployment a
SuperAdmin grant moved REAL money with NO durable cloud-side before/after record.
Decision: FAIL-CLOSED. Grants are a real-money op and the audit store is always
present in any real deployment — audit_serve.go REQUIRES a persistent data dir
for the trail (hard boot error otherwise), and a nil store arises ONLY from the
explicit CLOUD_AUDIT_DISABLED dev opt-out. So ApplyGrant now refuses the grant
with 503 BEFORE any money moves when AuditStore == nil: no unaudited money move.
The check sits after org validation and before the deposit, so the existing
validation errors (amount/cap/unknown-org) are unaffected.
Residual (documented, not regressed): a post-deposit Append failure (store
present but the write errors) keeps the money-moved success response — the money
DID land, and failing it would both misreport the grant and, absent idempotency,
invite a double-credit retry. It stays a loud log, backstopped by the
request-level AuditTrail middleware which independently records the request and
fails the request CLOSED on its own write error (serve.go / audit_middleware.go).
Finding 4 (deposit double-credit on retry — WIRED). commerce.Deposit posted
/v1/billing/deposit with no idempotency key, so a commit-then-timeout (cloud's
15s client) could drive an operator double-credit on retry. The commerce backend
DOES enforce idempotency: POST /v1/billing/deposit reads X-Idempotency-Key,
scoped billing-deposit:<subject> — a completed key REPLAYS the receipt, an
in-flight key 409s, an absent key is legitimately additive (never dedup by
amount). Verified in ~/work/hanzo/commerce/api/billing/deposit.go.
Fix: thread a DETERMINISTIC key from the grant. Deposit/post gained an
idempotencyKey arg sent as X-Idempotency-Key. ApplyGrant derives it as
sha256(org|amountCents|currency|source|nonce) where nonce is the operator-supplied
Idempotency-Key — so a retried grant carrying the SAME nonce dedupes at commerce
while two DISTINCT grants (even same org+amount) never collide, and a nonce reused
for a DIFFERENT amount still lands (dedup can never silently DROP a real grant).
No nonce => no key => additive default preserved; we do NOT fabricate a
content-only key (it would wrongly dedupe two legitimate identical comps).
Residual (cross-service, frontend): effective end-to-end once the operator
console sends an Idempotency-Key per grant attempt, reused verbatim on retry.
Finding 2 (GuardScoped over-visibility — NEEDS CROSS-SERVICE, not fixed here).
GuardScoped admits any validated principal with non-empty c.User()+c.Org().
SanitizeIdentity (middleware_identity.go) mints X-User-Id for EVERY validated org
member and mints the admin signal X-User-IsAdmin ONLY for a GLOBAL admin
(owner==adminOrg) — it DISCARDS the JWT org-level isAdmin bit for non-admin-org
principals. So a regular member's sanitized identity is byte-identical to an
org-admin's, and GuardScoped admits regular members to their own org's admin
panels (same-tenant over-visibility of financials + user directory; NOT
cross-tenant — ResolveScope pins to c.Org()). The finding is REAL, but the
prescribed fix (require org-admin) CANNOT be done in clients/admin/*: the
org-admin signal is not minted, and middleware_identity.go is read-only in this
scope. Requiring the only admin signal we have (c.IsAdmin(), GLOBAL) would break
the deliberate, test-locked org-scoped tier (scope_test.go admits org-admins).
Correct cross-service fix: in SanitizeIdentity's `case owner != ""` branch, mint
X-User-IsOrgAdmin=true when claims.IsAdmin (owner safe, non-KMS), add it to
authorityHeaders (stripped on ingress so unforgeable), then GuardScoped requires
c.IsAdmin() || X-User-IsOrgAdmin=="true". Left for the middleware owner.
Tests: TestGrantCredit_NilAuditStoreFailsClosed (503, no deposit, balance
unchanged); TestGrantCredit_IdempotencyKeyForwarded (same nonce=>same key,
distinct nonce=>distinct key, no nonce=>no key). Existing money must-pass tests
(TestGrantCredit_DepositLandsAndAudited, TestSuspendReactivate_*, TestAdminAudit_*)
still pass unchanged.
Finding 1 (undercount masked as healthy). core.OrgMoney swallowed the
Commerce.Spend/Credits errors to a silent zero, and /overview derived the
"commerce" source freshness from a SINGLE probe org — so if commerce was down
for 40 of 41 orgs the fleet spend/credits totals were an undercount while the
source still read "healthy". revenue.go / finance.go already report this
correctly via a `partial` flag + the core.ErrPartialRevenue sentinel.
Fix (decomplect to ONE partial pattern, not a second one):
- OrgMoney now returns (spend, credits int64, ok bool); ok is false when the
spend OR credits read failed — the SAME (row, ok) contract revenue.revenueOf
uses. An unwired commerce is NOT a failure (Spend/Credits return (0,nil) when
unconfigured), so ok stays true and Ready() still distinguishes not-configured.
- overview folds ANY per-org OrgMoney failure into the commerce source as
degraded, reusing core.ErrPartialRevenue, and derives freshness from the SAME
per-org reads the totals fold instead of a single probe org.
orgs is a per-ROW panel (OrgRow[] via OKList; NO sources[] channel): a failed
read degrades THAT org's row to an honest zero — there is no aggregate total to
mislabel, so no change beyond the signature. The customer per-row/detail call
sites are best-effort and unchanged in behavior.
Test: TestOverview_CommercePartialOnPerOrgError — commerce 500s for one org of
two; the commerce source reports not-ok with an error while the healthy org's
spend still contributes (honest partial total, not a hard panel fail).
The subpackage contained only replication_test.go referencing an undefined Producer
(no NewProducer/BackupOnce/etc. anywhere), so it never compiled and failed go vet.
Zero importers. Pike: the best code is no code.
Rewrite the top-level admin package as the Mount + aggregator only. state ->
core.State (exported fields), and every top-level handler + helper is retyped to
*cloud.Service[core.State] and calls core.* for the shared kernel:
- routes() registers the org-scoped panels (me/overview/orgs/users/usage/
analytics/bases) behind core.GuardScoped and the platform reads
(roles/applications/products/compute/o11y/sync + flags/waitlist) behind
core.Guard, then delegates to audit/customer/revenue/finance Routes().
- analytics.go keeps only the analytics-specific derivation (growth/retention/
churn/active/LTV) folding over the core activity model + spend series.
- o11y/compute/bases/waitlist/flags/types retyped to core.State + core.*.
Delete the now-moved audit.go/customers.go/grants.go/revenue.go/finance.go/
scope.go (their handlers live in the domain packages; their shared helpers in
core). doTokenFromEnv moves to admin config; iamAuditQuery moves to the audit
package.
Move tests with their code: audit store tests -> audit package; grantTag test ->
core; finance pure-math tests -> finance package. The full-mount integration
tests (admin/scope/cockpit/finance-aggregation) stay in the admin package and
now drive the real routes() so the harness mirrors Mount exactly. Behavior,
routes, tenant scoping and every assertion unchanged.
Introduce clients/admin/core as the subsystem's shared kernel — the State
struct (upstream clients + adminOrg + audit store) and the one-copy business
primitives every admin surface composes: the two-tier gate (Guard/GuardScoped),
the /v1 envelope writers (OK/OKList/OKRaw/Fail), CallerCreds, the tenant-scope
predicate (TenantScope/ResolveScope/ScopedOrgs/Descendants), the IAM fan-in
(ListOrgs/OrgMoney/FindOrg/Display/SrcOf/SourceStatus), the ONE credit-write path
(ApplyGrant + EmitAudit + grantTag/grantNote/CreditRequest) and the fleet
activity/time-series model shared by analytics and revenue
(CustActivity/TxnPoint/SeriesPoint/FleetActivity/SpendSeries + bucket helpers).
Carve one package per handler domain over that kernel:
- audit -> /v1/admin/audit{,/verify} (store-backed + IAM fallback)
- customer -> /v1/admin/customers* + /v1/admin/grants (list/detail/credit/
suspend/reactivate/grants ledger)
- revenue -> /v1/admin/revenue
- finance -> /v1/admin/finance (+ the pure ComputeFinance derivation)
Each domain imports core for the kernel and shared logic; no business logic is
duplicated. Routes are registered per-domain via <domain>.Routes(app, s).
Completes the bots surface (launch -> launch+list+stop). Both are thin,
org-scoped proxies onto the in-cluster bot-gateway (BOT_GATEWAY_URL, the same
server-side knob clients/bot uses), carrying the caller's validated tenant
context (X-Org-Id pinned to principal.Org, never a request param).
- GET /v1/bots normalizes the gateway's session rows into
{runId,task,surface,status,sessionUrl,startedAt}, deriving sessionUrl here
(the one place a session URL is built). Honest-empty {"bots":[]} on an
unconfigured/unreachable gateway, a non-2xx, or an undecodable body -- never 5xx.
- POST /v1/bots/:runId/stop returns {runId,status:stopped}; a run the caller's
org does not own is 404; an unreachable gateway is a clean 502.
Both require a validated principal so org-scoping can't ride a forged X-Org-Id.
Hermetic list+stop tests against a fake gateway assert normalization, sessionUrl
derivation, caller-org scoping, honest-empty, and 404/502 paths.
v1.805.5 adds the controller-side balance-gate fail-open (BALANCE_GATE_FAIL_OPEN_ON_ERROR)
on top of v1.805.4's nil-guard, so a broken/misconfigured Commerce billing backend
degrades to allowed-but-ungated instead of 500-ing every authenticated chat for
non-exempt users. Together they let the CR drop the commerce.hanzo.svc:8001 bridge
and keep commerceEndpoint unset per the in-process design.
ai v1.805.4 guards the nil balanceGate deref in resolveBillingKey that returned a
bare HTTP 500 on EVERY authenticated /v1 request (models, chat/completions,
messages) whenever commerceEndpoint is unconfigured (balance enforcement disabled).
RateLimitFilter calls resolveBillingKey before BalanceGateFilter's own nil guard,
so the embedded AI subsystem crashed every authed request while anonymous requests
(no token) 401'd correctly. Fail-open: a disabled billing subsystem resolves no
billing subject and never crashes a request.
Cloud edge for Hanzo Sentry (the /v1/sentry product face served by the embedded
o11y runtime):
- mountSentry registers the /v1/sentry/* wildcard, forwarding to the SAME gated
runtime handler the /v1/o11y wildcard uses (one runtime, two path families; no
path rewrite — the Sentry routes are literal /v1/sentry/... in the runtime).
- gate() now exempts the DSN-authenticated Sentry ingest writes (isSentryIngestPath:
POST /v1/sentry/{project}/envelope|store/, tight method+prefix+suffix match) from
the principal gate — the runtime authenticates the DSN key, not a Hanzo session —
while EVERY Sentry read/write API stays principal-gated (no cross-tenant leak).
Uses only the existing o11y runtime-handler API, so it builds against the current
pinned o11y (v1.5.12) and is inert (404) until the o11y dep is bumped to a build
that carries the /v1/sentry routes.
FOLLOW-ONS (coordinated separately):
- Bump the hanzoai/o11y dep to the tag containing the /v1/sentry routes to activate.
- Gateway needs a byte-identical isErrorIngestPath sibling for POST
/v1/sentry/{project}/envelope|store/ so the tokenless DSN ingest is not 403'd at
the edge (do NOT touch ~/work/hanzo/gateway here).
Test: TestGateExemptsSentryIngestButGatesReads (ingest exempt, reads/writes gated).
The Account.Token comment claimed the store is SQLCipher-encrypted and 'keeps
it encrypted at rest'. False: openStore uses cek.Open, which today runs the
no-key plaintext fallback (real WithRawKey-from-KMS lands with the connect
flow). Token column is empty today, so no plaintext secret ships.
Finishes the batch-F rework (the extraction landed in 850d4c0; this is the second,
first-principles half that the earlier cherry-pick stopped short of):
- clients/admin/money — type Cents int64; the unit lives in the type, so
ConsumedCents/MRRCents/CreditsCents collapse to Consumed/MRR/Credits. int64 casts
only at the operator-contract boundary (wire unchanged, JSON tags kept).
- clients/admin/digitalocean — do.go extracted from the inline client; Client.{Ready,
Balance,History}. Named digitalocean (not do) so it never collides with the local do var.
- clients/admin/commerce — reworked to the money.Cents unit + collapsed the always-equal
(org,user) into one subject; dropped the MRRCents + duplicate-rollup shims.
Applied cleanly (0 conflicts) atop the extraction + tenant->org main. Build ./... green,
vet clean, admin+commerce+digitalocean tests pass (pure-Go cek gate).
Extracts admin's inline upstream reader clients into self-contained, package-namespaced
units and gives billing a single value type. Decomplect + package-as-namespace, per the
Hickey/DRY bar:
- clients/admin/money — type Cents int64; the unit lives in the type (ConsumedCents/
MRRCents/… collapse to Consumed/MRR). int64 casts only at the operator-contract edge.
- clients/admin/commerce — Client.{Ready,Spend,Credits,Plan,Ledger,Costs,Deposit}
(Deposit = the one write). Collapses the always-equal (org,user) into one subject;
the bare-slug X-Org-Id+user invariant is baked into the client. Drops MRRCents +
duplicate rollup-balance shims.
- clients/admin/digitalocean — Client.{Ready,Balance,History}
- clients/admin/health — Client.{Ready,Up}
Wire unchanged (JSON tags kept, DTOs stay int64 cents). Integrated onto current main:
preserves the commerceinproc.BaseURL(...) in-process routing and the tenant->org
vocabulary. Build ./... green, vet clean, admin+commerce+integrations tests pass
(pure-Go, cek gate). Completes the batch-F rework the user authorized.
Fold the live social stack's publish + schedule onto the native /v1/social domain:
- Publish edge (publish.go): the ONE publishPost path — claim (at-most-once across
the HTTP handler and a scheduler tick), fan out to the channel's connected accounts
through the injectable Publisher seam, record the honest outcome (published + external
id, or failed + reason) on the post. POST /v1/social/posts/:id/publish + on-create
fanout (scheduled-for-now-or-earlier) both call it.
- Scheduler (scheduler.go): an in-process periodic sweep (the native equivalent of the
live stack's Temporal timer + hourly missing-post poller) that advances every org's
scheduled -> published when the time arrives; idempotent, per-org, clean shutdown.
Mirrors clients/commerce/sweep.go.
- Provider seam: fail-closed default (notConfiguredPublisher) that reports EXACTLY which
OAuth-app credentials are missing (the live orchestrator's env var names) and NEVER
fakes success. No Hanzo deployment carries provider creds today, so this is prod's
honest state. GET /v1/social/providers reports per-network publish-readiness.
- Store: token on accounts (SQLCipher-encrypted at rest; live stack stores it plaintext),
external_id/account_id/error on posts, ClaimForPublish/MarkPublished/MarkFailed,
DueScheduled (the ONE deliberate cross-org system read), RecoverStuckPublishing.
Tenant isolation preserved: every publish is org-scoped; the scheduler's cross-org sweep
only reads (org,id) to dispatch into the org-scoped path. 12 tests (9 new) prove publish
success/not-configured/no-account, idempotency, per-org isolation, the scheduler sweep,
crash recovery, and live capabilities. go build -tags 'cloud cloud_mount' ./... green;
frozen wire-order guard green.
Rebased blue's RED-reviewed encrypt-at-rest onto the CURRENT live-prod commit
(v1.786.185). The release train added stores blue's base never saw, all opening
PLAINTEXT — closed every one through the same cek.Open seam:
- tenantdb.go (the SOLE per-tenant opener → code/git/functions/tracker)
- gojabase, gatewaypolicy (gateway.db), dataroom (link_index.db)
Result: ZERO plaintext sql.Open("sqlite") left in the cloud data plane
(commerce self-encrypts under its own KMS key; cek internals excepted).
Flattened internal/cek → top-level one-word package cek. Trimmed the AI-slop
import comments to one line.
Verified CGO=1+SQLCipher: go build ./... green, cek tests green, and blue's
cek.Open shipping path pre-proven on ALL 47 real prod DBs (ciphertext + plaintext
shredded + exact row-count parity, 0 fail).
The frozen-fixture test ran in Go CI but not inside the image under the pinned
Alpine libsqlcipher, so a sqlcipher-dev pin/base bump that changes the on-disk
format could go green in CI while bricking prod. Add the frozen-format gate to
the Docker RED gate (beside TestEncryptionProof), and make requireCipher honor
SQLITE_REQUIRE_CODEC=1 (a would-be skip becomes a FAILURE) so the in-image gate
is airtight: any format change now fails the IMAGE build, not just Go CI.
Verified: SQLITE_REQUIRE_CODEC=1 go test -run TestFrozenFixtureOpens ./internal/cek = ok.
RED re-review closers:
1. Cross-version brick guard (vector b): the version-freeze was soft (unpinned
apk sqlcipher-dev resolves from the live Alpine repo). Now:
- Dockerfile pins sqlcipher-dev=4.6.1-r0 → a repo bump fails the build LOUDLY
(never a silent prod brick).
- Commit a FROZEN encrypted fixture (internal/cek/testdata/frozen/store.db
+ .dek, written under cipher_compatibility 4) + TestFrozenFixtureOpens that
copies it to temp, opens via cek, and reads a known canary row. A future
libsqlcipher format change makes the FROZEN store fail to open → red CI,
which build-time SQLITE_REQUIRE_CODEC (fresh-db only) cannot catch.
TestGenerateFrozenFixture (CEK_GEN_FIXTURE=1) regenerates it on an
intentional format rev.
2. Scope doc (vector c): cek package doc now states it provides CONFIDENTIALITY
at rest, NOT integrity/authenticity/anti-rollback vs a PV-write
(node-compromise) adversary — out of the stated read-only model. The
logical-id+epoch binding is deliberately NOT added (out-of-model complexity).
9/9 cek tests green on real SQLCipher (cgo) + nocgo; gofmt/vet clean. Migration
still deferred; live cutover gated on the user's supervised go.
RED review fixes on the encrypt-at-rest codec:
1. [HIGH] Shred the pre-migration <db>.plain.bak after the verified keyed reopen
(overwrite+remove, single call site in openEncrypted). Nothing reaped it
before, so every migrated store left a COMPLETE plaintext replica on the
volume forever — the exact threat the codec exists to kill. Test asserts the
backup is gone post-migration.
2. [MED] KEK no longer binds to the CLOUD_DATA_DIR-relative path (brittle: a
dir change bricked the plane). A random per-file id is stored in the sidecar
head (fileID(16) || wrapped-DEK) and the KEK derives from it — intrinsic to
the file, config-independent. Test moves a store to a new dir + wrong
CLOUD_DATA_DIR and it still opens.
3. [MED] Missing key on an encryption-capable build is now FATAL (fail-closed,
same posture as KMS), not a silent plaintext downgrade. Encrypting() is wired
into the boot log (serve.go). No new gate/env — keyed off the existing
CLOUD_KMS_MASTER_KEY_REF + build capability.
4. [MED] Parity gate gains a rowid-independent per-table content hash (commutative
multiset sum of per-row hashes over user columns), catching value mutations
that count+schema+integrity_check miss while ignoring benign rowid renumbering.
Dropped the inaccurate 'byte-faithful' wording. Cross-version cipher stability
is enforced by FREEZING libsqlcipher in the build (an at-open cipher_compat pin
is infeasible with mattn's URI-key requirement — the URI param is ignored and a
post-open pragma runs after mattn reads the header; proven); any mismatch fails
closed (test: corrupt/missing sidecar refuses).
5. [LOW] recoverInterrupted scrubs stale -wal/-shm; migration removes them before
the swap so a plaintext WAL never sits beside an encrypted db. No statement
logging on keyed conns; the key never rides a logged DSN/error.
8/8 tests green on real SQLCipher 3.53.1 (cgo) + nocgo fatal-key path.
The CGO+libsqlcipher cloud binary shipped encryption DORMANT: all 29 stores
opened via bare sql.Open("sqlite", path), so /var/lib/cloud/*.db (crm PII,
treasury ledger, wallets, audit, team/entitlements) were plaintext despite
CLOUD_KMS_MASTER_KEY_REF being present. The SQLCipher primitives were linked
(hanzoai/sqlite cek.go) but never USED.
internal/cek.Open is the ONE encryption-at-rest seam every store now routes
through. When the master key is configured it transparently, under a per-file
flock: mints a per-DB random DEK (SQLCipher page key), wraps it AES-256-GCM
under KEK=HKDF-SHA256(master, principal) in a <db>.dek sidecar, and — for an
existing PLAINTEXT file — migrates it to SQLCipher via sqlcipher_export, then
verifies schema + per-table row-count parity + integrity_check by re-opening
the encrypted copy EXACTLY as the app will, BEFORE an atomic swap. Fail-secure:
key set on a non-encrypting build is a hard error; unverifiable migration leaves
plaintext intact; wrong master fails closed; crash mid-swap is recovered.
Proven with real SQLCipher (TDD): plaintext→ciphertext header, row parity,
.dek unwrap, .plain.bak preserved, idempotent reopen, wrong-key rejected,
dev-mode gated, nocgo fail-closed.
Migration is deferred to a maintenance-window cutover (this image built via
arcd) — NEVER encrypt against the current binary, which cannot open SQLCipher.
Rewires 29 open-sites; base/o11y external-module DBs are a follow-on.
zip #5 deprecated (*App).Mount(prefix, h): it is exactly
app.All(prefix+"/*", zip.AdaptNetHTTP(h))
kept only as a behaviour-identical alias. Move every foreign-http.Handler
mount onto the explicit primitive so there is ONE way to put a route on the
app, and so the cloud money binary keeps building once zip deletes App.Mount.
Migrated all THREE route-mount call sites (the task scoped two; commerce is a
third — its exclusion note referred to the commerce.Mount *function*, not the
app.Mount(p, handler) inside it):
- clients/plugin/plugin.go:125 app.Mount(p.Prefix, h) → app.All(p.Prefix+"/*", zip.AdaptNetHTTP(h))
- clients/iam/iam.go:159 app.Mount(p, handler) → app.All(p+"/*", zip.AdaptNetHTTP(handler))
- clients/commerce/mount.go:147 app.Mount(p, handler) → app.All(p+"/*", zip.AdaptNetHTTP(handler))
Behaviour-identical (Mount IS this composition). Also updated the four doc
comments that named the deprecated zip.App.Mount so no dangling reference to a
soon-deleted method remains. grep-confirmed ZERO app.Mount( route-mounts left.
go.sum: removed the two inert zap-proto/zip v1.3.0 hash lines (nothing in the
module graph requires v1.3.0 — orphan cruft a working `go mod tidy` would drop).
Full `go mod tidy` is blocked by a PRE-EXISTING, unrelated force-moved tag on
luxfi/keys@v1.2.2 (server serves h1:nuD+y5…; committed go.sum records
h1:XH5mRm…), so the two dead lines were removed surgically instead — luxfi/keys
and every other entry left byte-identical to origin/main. No GONOSUMCHECK hack.
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Per the one-authority directive: commerce must not require its OWN Organization
record; IAM is the sole org/user/auth authority. The service-token path already
auto-projects the org from X-Org-Id via the cached GetOrCreate resolver, but the
IAM-principal path c.Next()'d without resolving the org — it depended on
iammiddleware running upstream, and when it hadn't, GetOrganization MustGet-
panicked (500) / the org was absent, so an IAM org with no pre-existing commerce
row could not view or create billing. ensureIAMOrg now resolves the validated
X-Org-Id through the SAME cached GetOrCreate resolver on the IAM path too
(idempotent), so any IAM org "just works" and a thin billing record is
auto-projected on first use — commerce derives the org from IAM, never its own
table. This is also the root cause behind the "hanzo has no commerce org record"
wall on the live 2-org proof.
The billing-account CRUD (List/Create/Get/Update/Delete + project bindings +
members + loadOwnedAccount) called middleware.GetOrganization, which MustGet-
panics (→ recovered 500) when the request's org has no commerce Organization
record — the same class as the already-deployed spend-alert fix. Switch all to
GetOrganizationOK with a safe default (reads → empty, mutations → 400/404), so
an IAM principal whose org lacks a commerce record gets a clean response, not a
500. Completes the metering-path panic hardening.
Bump github.com/zap-proto/zip v1.3.0 → v1.5.0 (verified drop-in) and move
subsystem teardown onto zip's OnShutdown hook, deleting the hand-rolled
ShutdownAll reverse-loop.
Before: serve.go called ShutdownAll(specs) — a reverse-mount teardown loop —
BEFORE app.ShutdownWithContext. Subsystems were torn down while the listener
was still accepting and in-flight requests were still draining: a latent race
(a request could use a store that teardown had just closed).
After: MountAll registers each enabled spec's ShutdownFunc via app.OnShutdown
right after the subsystem mounts. zip drains those hooks LIFO — AFTER the
listeners stop accepting and in-flight requests drain — so registration-at-mount
reproduces the exact reverse-mount teardown order ShutdownAll gave, minus the
race. app.ShutdownWithContext now owns the whole teardown.
Scope is deliberately narrow: MountSpec, MountAll, Wire(), Typed, and the
OwnsHealth health loop all STAY (the imperative composition-root flatten was
proven unsound — a generic /v1/:name/health route is shadowed by 3-segment
subsystem routes). Only ShutdownAll and its serve.go call site are deleted;
audit/gateway-policy/telemetry teardown keep their positions.
Test: build_onshutdown_test.go drives a real in-flight request over a loopback
listener, shuts down mid-request, and proves (1) Shutdown blocks until the
request drains and (2) the MountAll-registered hooks then run LIFO = reverse
mount order. A second test locks the enablement axis + nil-Shutdown guard.
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The Sentry error-ingest wire endpoints (POST /v1/o11y/api/<project>/envelope|store/)
authenticate with a DSN public key downstream in the o11y handler, not a Hanzo
principal. cloud's o11y gate() 403'd them for lacking X-User-Id, so the tokenless
ingest that the gateway allowlists (isErrorIngestPath) was blocked one layer deeper
— error-tracking could never ingest end-to-end.
Add the cloud-side counterpart: gate() now also exempts isErrorIngestPath
(byte-for-byte the gateway's matcher — method + /v1/o11y/api/ prefix + envelope|store
suffix, never a bare prefix). Reads under /v1/o11y/api/vN/... and the Issues
list/detail/update stay principal-gated; the exempted ingest still fails closed on a
bad/absent DSN key (401/503). Tested: TestGateExemptsErrorIngestButGatesReads.
The live signed-webhook e2e (#118) got past HMAC verification and then
500'd: resolveWebhookOrg called middleware.GetOrganization — gin MustGet
— but webhook ingress runs OUTSIDE the auth-token group, so no
middleware ever set "organization" and every signature-VALID provider
delivery panicked ("key organization does not exist"). Switch to the
GetOrganizationOK variant that exists precisely for signature-verified
sessionless ingress; the header/env/default fallback chain below it now
actually runs.
Regression test drives resolveWebhookOrg on a bare sessionless gin
context and asserts no panic.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Every inbound Square webhook 401'd 'square processor not configured':
payment/providers/square registers an EMPTY Provider whose Configure()
only runs from BD's per-tenant charge resolver — but tryValidateWebhook
(the sessionless inbound path) reaches the registry slot directly. The
env-registering thirdparty/square init could never win the slot either
(import-order clobber / defer-to-existing), so webhook validation had no
configured processor in ANY deployment — 100% of live Square deliveries
were rejected.
init() now seeds Configure() from the deployment env (same vars +
SQUARE_ENVIRONMENT sandbox switch thirdparty reads; charge creds required,
webhook fields optional with ValidateWebhook's own fail-closed contract).
BD's per-tenant Configure still overrides per request for outbound money.
Regression guard drives env -> configFromEnv -> Configure ->
ValidateWebhook over a real Square-spec HMAC (url+body), plus tamper and
sandbox-switch cases.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Standalone commerce was swept by a k8s CronJob (curl image + cluster hop +
token secret) POSTing /v1/billing/auto-recharge/run-all every 15m at
commerce.hanzo.svc:8001. With commerce folded into the cloud binary that
loop is redundant moving parts: the binary now dispatches the SAME request
loopback through its own embedded gin handler — full middleware chain
(TokenRequired service-token branch -> PlatformOnly mint gate -> datastore/
KMS context), byte-identical to the wire path (guarded by
TestSweepOnceWireContract).
- interval: COMMERCE_AUTORECHARGE_INTERVAL (default 15m; 0/off disables;
garbage disables fail-safe — a card-charging loop must never guess)
- token: COMMERCE_SERVICE_TOKEN (absent -> disabled with a loud warn,
not a 403-every-tick loop)
- first fire after one FULL interval (never front-run the outgoing CronJob
during rollout)
- runs only where commerce mounts (single-writer pod; reader role never
mounts commerce) -> exactly one sweeper, same guarantee the single
CronJob schedule gave
- stops with Embedded.Stop
Unblocks deleting the commerce-auto-recharge CronJob + the standalone
commerce/commerce-sandbox pods (#118).
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Turns the agentic-marketing content loop from scaffolding into a working
edge pair, then closes every finding from the adversarial review.
Edges (wired at Mount, fail-closed until configured):
- Generator: zen5 copy via deps.AI.ChatCompletion (metered Bill.Gate/MeterUsage)
+ studio assets via the ComfyUI Qwen-Image-Edit-2511 graph.
- Distributor: hanzoai/social Public API fan-out, per-brand key custodied in KMS.
errNotConfigured until a brand connects a key — no key, no fan-out.
Hardening:
- Publish TOCTOU interlock: a per-item, store-backed lease (framework fw_locks,
the ONE cross-process coordination primitive) serializes concurrent publishes
of the same item across drivers and pods; non-idempotent fan-out runs once.
- source_media SSRF gate: the URL validator fails closed on unparseable/backslash
input.
- external_ids is server-managed: the before_save guard rejects a client write;
Publish records it under the lease as a trusted server write.
Adversarial regression suite pins each guarantee from the outside (raw framework
PUT, direct ops, concurrency): red_adversarial / red_rereview / red_final /
red_lease_final, plus blue_hardening and per-file unit tests.
One open item recorded as a HARD GATE in clients/content/LLM.md: lease
TTL-preemption re-opens the double-post window for a brand with ~15+ slow
channels (fan-out (N+1)x20s > 5m TTL, external_ids recorded only at fan-out end).
ZERO exposure until a brand connects a social key; the fix (lease heartbeat/renew
preferred) MUST land in the same change that connects the first real key.
Claude-Session: https://claude.ai/code/session_013jh8aka8q8RvhhVQ1psMeW
Two INFO items from RED's #250 re-review:
- TestFrameworkContentModulesLinked asserted only the module registry. erp's
ledger-posting HOOKS register in a SEPARATE init() step, so a future split of
registerHooks() out of erp's module init() could drop the hooks while the guard
stayed green. Add framework.RegisteredHookCount() and assert it > 0 so the guard
fails if erp's hooks (computeJournalTotals, journalEntry/paymentEntry
submit+cancel, …) are ever unlinked from the binary.
- gojabase Config.DataDir comment said "{tenantSlug}.db" (the pre-C1 name); the
on-disk segment is TenantSegment (injective, traversal-safe base32 of raw org
bytes). Comment now matches the code.
go build + go test ./subsystems/... ./clients/framework/... green.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Net-new per-org ad-campaign domain on the ONE cloud framework (zip/Fiber +
cloud.Deps + per-org SQLite), the same shape clients/crm uses — twin of crm.
Registered in subsystems.Wire() right after crm; frozen wire sequence updated.
/v1/ads surface (all org-scoped, tenant-isolated on the bearer owner claim):
GET /v1/ads/health (auto, serve.go liveness)
GET /v1/ads/summary per-org roll-up (total/active/budget/spend)
GET /v1/ads/campaigns list (?status=)
POST /v1/ads/campaigns create
GET /v1/ads/campaigns/:id detail
PUT /v1/ads/campaigns/:id update
DELETE /v1/ads/campaigns/:id delete
Campaign = root of the ad hierarchy (campaign -> ad sets -> ads; ad-set/ad legs
hang off this seam): Platform (meta/google/tiktok/x), Status (draft/active/
paused/completed), Objective, Budget/Spend (cents).
Tests: per-org isolation + full CRUD + summary round-trip, both green.
In-process fold of github.com/hanzoai/social (social-backend/frontend/
orchestrator, a Postiz-style scheduler) onto the ONE cloud framework
(zip/Fiber + per-org SQLite), twin of clients/crm + sibling of the
marketing fold. NOT a proxy to the standalone social pods.
Two entities faithful to the live Public API (clients/content publish.go
already talks to it): Account = a connected channel (the stack's
integration), Post = content published/scheduled to a channel. Scheduling
is a Post with status=scheduled + a future scheduleAt, not a third entity.
Every query filters WHERE org=? on principal.Org (validated bearer owner,
HIP-0026); one tenant can never read/mutate another's rows. Registered in
subsystems Wire() after crm with a ctxShutdown that closes the DB.
Surface (org-scoped, /v1 only): summary + accounts CRUD + posts CRUD.
Generic liveness serves GET /v1/social/health.
Build: go build -tags 'cloud cloud_mount' ./... GREEN; 3 store tests pass
(per-org isolation across accounts+posts, post CRUD+summary, account CRUD).
Mount github.com/hanzoai/marketing in-process on the ONE cloud framework
(zip/Fiber + cloud.Deps + per-org SQLite), the same shape clients/crm uses —
NOT a proxy to a standalone marketing pod. Registered in subsystems.Wire()
right after its twin crm; frozen wire sequence updated to match.
/v1/marketing surface (all org-scoped, tenant-isolated on the bearer owner claim):
GET /v1/marketing/health (auto, serve.go liveness)
GET /v1/marketing/summary per-org roll-up (total/active/budget/spend)
GET /v1/marketing/campaigns list (?status=)
POST /v1/marketing/campaigns create
GET /v1/marketing/campaigns/:id detail
PUT /v1/marketing/campaigns/:id update
DELETE /v1/marketing/campaigns/:id delete
Campaign faithful to the repo domain: Channel (email/sms/social/meta/google/
tiktok), Status (draft/active/paused/completed), Objective, Budget/Spend (cents).
Genetic-optimizer / ML-forecasting / ad-platform integrations not folded yet.
Tests: per-org isolation + full CRUD + summary round-trip, both green.
Bumps github.com/hanzoai/ai v1.805.2 -> v1.805.3. v1.805.3 pins the GenAI
tracer at adopt time so the embedded o11y/SigNoz runtime reassigning the
process-global OTel tracer provider can no longer redirect gen_ai spans off
the o11y sink — they reach o11y_traces. No cloud source change; go.mod/go.sum only.
Add clients/content — the ONE Go-native replacement for the bespoke karma
Python pipeline. A framework app-lane (module "marketing": Campaign, SocialPost,
Asset) + a thin /v1/content/* control-plane. Store-less: content IS framework
documents; the subsystem is a stateless orchestrator over the framework store +
the zen5/studio/social edges.
- lifecycle.go: ONE state machine (draft→in_review→approved→queued→published,
+archived), a pure value read by the before_save hook (enforces edge legality
at the storage boundary) and the transition endpoint (same check + fan-out).
- hooks.go: before_save gate on every publishable DocType.
- content.go: board (cross-DocType aggregate), lifecycle, transition (+best-effort
distribution), generate/publish/channels. Org-scoped via principal.Org; never a
5xx from a foreseeable condition (fail-closed 503 / honest 4xx).
- generate.go/publish.go: Generator + Distributor seams with fail-closed defaults;
real zen5 (deps.AI) + studio (ComfyUI) + hanzoai/social wiring documented for the
follow-up. Exported Generate/Publish/Transition are the ONE impl the HTTP surface
AND the automations connector call.
- framework: add Get (read-one), exported ErrNotFound/ErrConflict/ErrBadRef +
IsValidationError so in-process callers classify errors without string-matching.
- automations: connector_content.go exposes content_generate/transition/publish as
flow steps + MCP tools (in-process, org-scoped) so the loop runs autonomously
(cron flow: generate → wait_for_approval → transition → publish).
- subsystems: Wire content after knowledge; frozen wire-order test updated.
Tests: lifecycle table, before_save hook, end-to-end loop over the real framework
store (install→create→board→transition→publish), forge-403, cross-org isolation.
go build ./... + go test (content/framework/automations/subsystems) all pass.
Claude-Session: https://claude.ai/code/session_013jh8aka8q8RvhhVQ1psMeW
#248 (Wire() composition root) rewrote subsystems.go and dropped the blank
imports for cms/erp/help. Those three are NOT mount subsystems — no HTTP surface,
never in Wire(). Each registers DocType fixtures and, for erp, ledger-posting
lifecycle hooks (computeJournalTotals, journalEntry submit/cancel, paymentEntry
submit/cancel, …) into the always-on clients/framework engine from a package
init() (framework.RegisterModule). Dropping the imports left them out of the
binary: /v1/framework/* carried no erp/cms/help and the erp ledger hooks were
silently gone. No mount test caught it (frozen[]/Wire() cover only mount specs).
Fix (RED-verified): re-add the three blank imports under an explicit "framework
content modules" comment; add framework.RegisteredModules() + a guard test
(TestFrameworkContentModulesLinked) asserting the engine carries each lane so the
drop cannot silently recur. Also relax TestDepGatedSubsystemsFailClosed to accept
any non-2xx — a dep-disabled subsystem denying 403 is fail-closed; >=500 was
wrongly strict (o11y denies 403, ai returns 5xx).
go list -deps ./cmd/cloud shows cms/erp/help; go build ./... green;
go test ./subsystems/... ./cmd/cloud/... -race green.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The S2S metering path (ScopeRules → ListSpendAlerts; AuthorizeVerdict →
AuthorizeSpendCap) can reach these handlers with no "organization" context key
when X-Org-Id does not resolve to a commerce org (e.g. the agents scheduler
probing an unprovisioned org). middleware.GetOrganization MustGet-panicked
there → a recovered 500 storm on the money path (observed live: "key
organization does not exist", ~every 5-15s). Switch the metering-verdict
handlers + billingSubject to GetOrganizationOK with a safe default — no org
means no org-scoped caps, so ListSpendAlerts returns empty and
AuthorizeSpendCap allows — instead of panicking. Regression test proves neither
handler panics with no org.
types.Claims gains Project + BillingAccount; principal.BillingAccount(c) reads
the gateway-minted X-Billing-Account-Id, added to subScopeHeaders so the raw
client copy is stripped on ingress and re-injected only for a validated
principal (mirrors X-Project-Id). It is an ATTRIBUTION hint only — the debited
account is always resolved server-side by commerce from the org's
ProjectBinding, never trusted from the header — so a mislabelled account can
only ever misattribute the caller's OWN spend within its own org, never
redirect spend to another tenant. Inert until IAM emits the billing_account
claim + the gateway mints the header (slices 63/64).
Also fixes a stale test left by de2a97a3 (/metrics is now a console page, not an
API 404 — dropped from the must-404 list).
Replace the init()-registry (blank imports + cloud.Register(name, order, …) +
magic order-ints) with ONE explicit subsystems.Wire() []cloud.MountSpec. Slice
position IS the mount order — MountAll iterates it as-given, ShutdownAll in reverse.
Core (build.go/serve.go/cmd): delete MountSpec.Order, Register,
RegisterWithShutdown, the HealthOwner option func, and var Registry. MountAll and
ShutdownAll take the []MountSpec slice; Serve threads it in (cloud never imports
subsystems → no cycle). cmd/cloud + cmd/hanzo call subsystems.Wire(). Typed and the
factory hooks (RegisterKMSClientFactory, RegisterCommerceClientFactory,
RegisterOrgScopeResolver, RegisterPushBuilder, RegisterTelemetryInstaller) are
untouched — a different mechanism.
In-repo (~64 clients/*): delete each init() registration; export the mount fn where
unexported (commerce.MountFromDeps, o11y.MountO11y) and the shutdowns. ctxShutdown
bridges the func() error shutdowns in one place.
Externals wired explicitly; the wave-2 tags no longer self-register:
ai v1.805.2 (@150 catch-all), o11y v1.5.12 (@70 wildcard) — origin/main already
pins these (genai-span) but did NOT wire them, so main currently DROPS ai + the
o11y wildcard from its live registry (main's own TestRegistryAssemblesSubsystems
fails on that). This composition root RESTORES both. authz v1.10.7 (@70),
licensing v0.1.3 (@110, Mount is already a MountFunc — wired direct), metrics
v1.110.2 (@40, takes its own metrics.Deps via mountMetrics). vfs unchanged: it
never registered a subsystem, so it is not wired.
Order is proven by subsystems.TestWireOrderMatchesFrozen against the sequence
captured empirically from origin/main @c504d2b (68 self-registering specs) with ai
+ the o11y wildcard restored at their order-int slots. The Service[state] refactor
reshuffled same-order tie positions on main; the frozen sequence adopts main's new
order (functionally inert — tied subsystems own disjoint route prefixes).
Also fixes a pre-existing clean-main build break (#247): clients/sign/sign.go
referenced an undefined `log` — main does not compile without it, so the whole tree
(subsystems → cmd/cloud) failed to build. One-line fix (log → deps.Logger); flag
for the owner to factor out if preferred.
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
#247's VFS-nil health-only guard was authored against the pre-refactor sign.go
(which had a local `log` var); it auto-merged cleanly onto the concurrently-landed
cloud.Service refactor (principal.Org / s.Log) but referenced an undefined `log`,
red-ing the build (clients/sign/sign.go: undefined: log). Point it at deps.Logger
(guaranteed non-nil by the guard above it). No behavior change.
go build ./... + go build -tags "libsqlite3 sqlite_fts5" ./cmd/cloud +
go test -race ./clients/{gojabase,captable,sign,dataroom}/... all GREEN.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Second wave of RED's consolidation review — the mediums that needed new
hanzoai/captable + hanzoai/sign bundle versions (now merged) plus their cloud-side
wiring and tests.
M5 (captable bundle @119bf9c0) — capTable summed ALL option rows regardless of
status, so EXERCISED/EXPIRED/CANCELLED grants inflated fullyDilutedShares and
every ownership %. Fixed to count OUTSTANDING options only. New
TestDilutiveOptionsExcludeTerminal proves terminal-state grants don't dilute.
M8 — ONE blob seam in gojabase (Config.Blob + a tenant-scoped globalThis.__blob =
{ put(key,b64), get(key) }), keys namespaced {Name}/{TenantSegment} so a bundle
can only reach its own tenant's objects. sign (bundle @315a3fe6) now routes PDF
bytes THROUGH it instead of inlining 32 MiB base64 in document_data.data: create
stores the original blob key, seal writes a sealed key, view/download read via
__blob.get. The sign leaf passes deps.VFS as the seam (health-only without it).
document_data holds only the key now (type BLOB_KEY). Same VFS/S3 plane dataroom
uses — one blob strategy, not two.
M7 (sign bundle @315a3fe6) — SEQUENTIAL signing order was declared but never
enforced. assertTurn gates signField+signComplete so a later signer can't act
until every earlier signer has SIGNED. New TestSequentialSigningOrder proves an
out-of-turn signer is refused (403) and the turn opens once the earlier one signs.
sign tests now inject an in-memory VFS and assert PDF bytes land on the seam
(tenant-scoped key), NOT in the tenant SQLite. All four gate packages pass under
go test -race against the published module versions.
Gate: go build ./... + go build -tags "libsqlite3 sqlite_fts5" ./cmd/cloud +
go test -race ./clients/{gojabase,captable,sign,dataroom}/... all GREEN.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
cloud installs the ONE global tracer provider wired to the o11y in-process trace
sink, but in in-process-sink mode it sets no OTLP/ZAP exporter endpoint. The embedded
hanzoai/ai module then found no endpoint and DISABLED its gen_ai span emit, so the
gen_ai plane in o11y_traces was dark (cloud's own request spans flowed; every LLM
call's gen_ai span was gated off). The provider install + adopt also lived ONLY in
cmd/cloud/main.go, leaving every 'hanzo <svc>' entrypoint (which shares cloud.Serve)
telemetry-dark.
Hoist the bootstrap into cloud.Serve — the ONE serve body cmd/cloud AND every
'hanzo <svc>' share — so both install the provider and adopt it into ai identically,
BEFORE MountAll mounts ai (ai's InitTelemetry reads the adopted-ready flag at mount).
cloud.Serve cannot import clients/o11y (clients/o11y imports cloud -> cycle), so the
concrete bootstrap (SetTracerProvider + aiobject.AdoptHostTracerProvider) lives in
clients/o11y and registers via cloud.RegisterTelemetryInstaller — the same cycle-free
inversion as RegisterKMSClientFactory / RegisterPushBuilder. Serve flushes the
provider on shutdown BEFORE ShutdownAll tears the sink down, so buffered spans drain.
Removes cmd/cloud/telemetry.go (moved to clients/o11y/traceprovider.go). Repins ai to
the branch carrying AdoptHostTracerProvider.
Red review (SHIP): MED-1 — this hoist (both entrypoints adopt identically). Tests:
cloud.installTelemetry dispatch/no-op (telemetry_test.go); clients/o11y
installTraceProvider adopt-latching + OTLP-env-retire + disabled-noop
(traceprovider_test.go).
ai: -> v1.805.1-0.20260710213829-b5eb789d3878 (feat/genai-span-host-provider @ b5eb789d)
* cloud: IAM-native vocabulary — org/user/project everywhere, tenant concept removed (TenantDB→OrgDB)
"tenant" was a non-IAM synonym for org. Replace it with the IAM-native nouns
org / user / project / billing account so there is ONE name for the concept.
Root framework primitives (package cloud) — now 100% tenant-free:
TenantDB -> OrgDB (tenantdb.go -> orgdb.go)
TenantStore[T] -> OrgStore[T]
NewTenantStore -> NewOrgStore
tenantDBPath / openTenantDB -> orgDBPath / openOrgDB
TenantScopeResolver -> OrgScopeResolver (tenant_scope.go -> org_scope.go)
RegisterTenantScopeResolver -> RegisterOrgScopeResolver
Dropped the legacy X-Tenant-Id / X-Tenant-ID entries from the identity
strip-list (nothing reads them; X-Org-Id is the live header).
Cross-subsystem seams:
principal.Tenant(c) -> principal.Org(c) (clients/principal; no collision —
Project/User kept, already IAM-native). ~50 call sites + ~70 stale doc refs.
types.TenantConfig -> types.OrgConfig; CommerceClient.GetTenantConfig ->
GetOrgConfig (+ commerce client impl, disabled/rpc/entitlements consumers,
the cloud.OrgConfig alias in deps.go).
Adopters deep-cleaned (identifiers, comments, test names): clients/code, git,
functions, tracker, provisioning, projects. Docs: subsystems.go composition
comments, README.md, and a new LLM.md "Framework doctrine" section.
GATED EXCEPTION (flagged, intentionally unchanged): clients/platform derives
LIVE k8s namespaces, registry image refs, and quota/limit objects from a
"tenant-<org>" string prefix. Renaming it orphans deployed namespaces + built
images, so the literal string is retained behind a // NAMING(gated) note in
clients/platform/k8s.go. The rbac test file + its funcs were renamed
(tenant_rbac_test.go -> org_rbac_test.go).
Out of scope (separate waves; independent domains, touched only for the
principal.Org call-site + stale-ref fix): commerce internal tenant tables
(active commerce-dissolve branch), runner tenantsource, treasury ledger, and
other subsystems' own tenant vocab.
Build: GOWORK=off CGO_ENABLED=0 go build ./... -> exit 0.
Tests: every touched package green. The 38 pre-existing failures (fakeAI
missing Embed; undefined Producer; commerce integration tests; o11y/graph/zt/
kms route-precedence & service tests) fail identically on the base commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* rebase: apply org vocabulary to post-branch main commits (metering/HA/tracing comments); drop unused TenantConfig alias
---------
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Follow-up to the reader-tier commit — align the CLOUD_WRITER_URL/
CLOUD_READER_RETRY_BUDGET/CLOUD_WRITER_LEASE struct-literal keys with gofmt so
the CI fmt gate is clean. No behavior change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
RED adversarial review of the one-binary consolidation (gojabase + captable/sign/
dataroom folds). Fixes the two ship-blockers plus the in-repo mediums/slop.
C1 (HIGH) — cross-tenant collapse via non-injective filename encoding.
gojabase.slugify ToLower+folded [^a-z0-9._-]→'_', collapsing DISTINCT owners
(Acme/acme, "a b"/"a_b") onto ONE per-tenant SQLite file — a cross-tenant break,
the exact fold principal.Tenant keeps the org verbatim to avoid. Replaced with
gojabase.TenantSegment: lowercased unpadded base32 of the RAW org bytes —
INJECTIVE (distinct orgs never share a file) and traversal-safe ([a-z2-7], never
"."/".."/"/"). Table test proves injectivity (Acme≠acme≠a b≠a_b) + traversal
containment. IAM owner charset permits case/separator variants, so this was a
real (not theoretical) collapse. Fresh-tenant-only: the folds pre-date any
release, so no on-disk data uses the old names — nothing to migrate.
H2 (HIGH) — code==docs on staging. captable/sign/dataroom were un-staged in
config.go (stagedSubsystems={iam,ingress}) yet every leaf/README claimed
"STAGED / standalone keeps authority". Evidence (do-sfo3-hanzo-k8s): NO
standalone captable or dataroom deployment exists; the esign pod runs but its
SQLite holds ZERO tenant data (User=0, DocumentData=0, Recipient=0, Field=0 —
only BackgroundJob/RateLimit churn). Decision (a): keep un-staged, delete the
false comments (captable.go, sign.go, dataroom.go, subsystems.go, README ×2).
No migration needed (dead/empty apps).
M3 drop dead commercesvc api.Route(/v1/{billing,checkout,store,subscription,
account}) — unreachable in embed mode (only /v1/commerce/* mounted, no
rewrite); money path is clients/billing, unaffected. Removed unused import.
M4 gojabase per-tenant *sql.DB cache is now an LRU (cap CLOUD_GOJABASE_MAX_DBS=256)
with idle eviction (CLOUD_GOJABASE_IDLE_TTL_SEC=300) and dispatch PINNING —
only idle handles evict, never a DB mid-transaction. Bounds fd/memory for N
tenants.
M6 dataroom deny_list is now ENFORCED in view.authenticate and WINS over allow
(checked first); surfaced on linkOut. One shared email matcher (DRY).
M9 sign refuses to boot with a self-signed cert in a production env — requires the
KMS-custodied CLOUD_SIGN_CERT_PEM/KEY_PEM; dev still self-signs.
M10 added a -race concurrent multi-tenant dispatch test (heavy eviction churn +
case-variant tenants) proving pool + cache + eviction are race-free and
isolation holds under concurrency.
Slop/decomplect:
- goja: a response with no explicit valid status now FAILS CLOSED (default 500,
rolls the transaction back) instead of committing at a defaulted 200.
- newID (gojabase) + randKey (dataroom): a crypto/rand failure now returns an
error / fails the dispatch instead of a predictable time/zero fallback.
- subsystems.go: removed 6 duplicate blank imports (pubsub, eval, exec, plan,
plugin, pricing).
- dataroom: unified on ONE tenant encoding — the object-store key prefix now uses
gojabase.TenantSegment, matching the SQLite filename (was raw org).
Gate: go build ./... + go build -tags "libsqlite3 sqlite_fts5" ./cmd/cloud +
go test -race ./clients/{gojabase,captable,sign,dataroom}/... all GREEN.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The metering decorator (metered_ai.go) wraps deps.AI so every chat + embed
call is authorized against the caller's org balance/budget/freeze BEFORE the
call and debited its billing account AFTER — the single DRY chokepoint through
which no inference runs unattributed. The billing scope (org, project) rides
the request value (ChatRequest/EmbedRequest), so it is compile-time impossible
to call AI without declaring who pays; no bypass, no side-channel key.
Closes the exempt holes — clients/code (index+search+ask), clients/knowledge
(search + index embeds), clients/crm (application screen) — all previously ran
the balance-exempt M2M path unmetered; each now threads org/project to the
chokepoint.
- types.AIClient.Embed takes *EmbedRequest (scope); ChatRequest carries
Org/Project; ChatResponse surfaces token counts for exact metering.
- httpAI stays pure transport but surfaces resp.Usage tokens and stamps
hanzo.org/hanzo.project on its gen_ai spans (feeds per-tenant o11y isolation).
- token-based micro-USD debit (CLOUD_AI_PRICE_UUSD_PER_1K, default $2/1M tok);
metering.Usage gains AmountMicros so a sub-cent call meters exactly instead
of rounding to zero and slipping through unbilled.
- reuses ResourceMeter over Deps.Metering — same per-org invariants; a
transparent pass-through when commerce is unconfigured (dev never blocked).
- fixes a latent slice-1 test break (crm/agents AIClient fakes lacked Embed).
- tests: pricing, scope forwarding, system-call, ModelLister preservation.
The console catch-all 404'd /metrics because it was in apiPrefixes (the Prometheus
scrape-path convention). But the real Prometheus surface is on the SEPARATE ops
listener (:9090, healthMux); on the product API (:8000) /metrics is the console
MetricsModule page. Drop /metrics from apiPrefixes so it falls through to the SPA
shell. One surface per port: :8000 product+SPA, :9090 ops. (ServiceMonitor repointed
to the ops port in the same change.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The cloud writer embeds an exclusive-lock ZapDB KMS store. A new probe
(clients/kms.TestConcurrentOpen_LiveWriterStoreIsNotROShareable) proves that
opening that store READ-ONLY while the writer is live FAILS ("Log truncate
required to run DB") — Badger's RO open replays the live memtable WAL and
refuses to truncate it. So the prior groundwork's assumption that a reader can
open the KMS store RO off the writer's PVC is false for a LIVE writer (only the
sequential close-then-reopen case worked). The audit SQLite store IS
concurrently shareable (audit/shareability_probe_test.go); the KMS store is the
one that is not, and every mutation is audited on the writer anyway.
Reader tier is therefore a transparent, always-ready reverse proxy (opens no
stores) to the single writer:
- reader_proxy.go: CLOUD_ROLE=reader boots serveReaderProxy BEFORE BuildDeps —
forwards every request to CLOUD_WRITER_URL, streams SSE, preserves inbound
Host. Dial-only retry (retryTransport) absorbs the writer's roll gap: it
retries ONLY when the connection was never established (no ready endpoint /
refused), so a non-idempotent POST is never double-executed; bounded by
CLOUD_READER_RETRY_BUDGET (default 25s) then 502.
- The reader Deployment rolls RollingUpdate(maxUnavailable:0), so the edge
Service always has a ready endpoint — this removes the ~30s console blip that
the writer's Recreate/replicas:1 causes today.
Writer zero-gap roll (opt-in, default OFF = byte-identical Recreate):
- writer_lease.go (+_unix/_other): CLOUD_WRITER_LEASE takes an exclusive fcntl
flock on {DataDir}/.writer.lock BEFORE opening the RWO stores and releases it
LAST at shutdown (after every store closes). A surge writer blocks until the
old one releases, so the exclusive ZapDB/audit stores are handed off, never
double-opened. Fail-closed on timeout.
Removes the dead ReaderGuard (the reader no longer runs the full pipeline; it is
the proxy). Unset CLOUD_ROLE + unset CLOUD_WRITER_LEASE ⇒ writer, byte-identical
to today.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
BillingAccount is a funding entity separate from the org: it holds a spend
limit + level + freeze and funds 1..N of the org's projects (dedicated or
shared). A project resolves to its account via ProjectBinding; usage debits
carry Transaction.AccountId so an account's balance and calendar-month spend
sum from the SAME append-only ledger the balance gate + per-scope caps read —
no parallel ledger, no stored total, never drifts.
The authorize verdict layers an account cap + freeze on top of the per-scope
caps (most-restrictive-wins, fail-closed). Resolution is ALWAYS server-side
from the org namespace — the account is never a forgeable client header.
Backward-compatible: (org,"default") folds to the org-wide pool, so an org
with no account behaves byte-for-byte as before.
- models/billingaccount, models/projectbinding (+ hashid kinds 284/285,
query reserved-kind list)
- Transaction.AccountId indexed ledger axis
- resolveAccountId + accountSpentCents + account cap/freeze in AuthorizeSpendCap
- real account + project-binding CRUD replacing the org-wrapper stubs
- tests: account cap, freeze kill-switch, dedicated-binding isolation,
backward-compat pool fold
zip v1.2.1 -> v1.3.0 rides zap-proto/fiber v3.2.1 (gofiber v3.2.0 + native
ServeMux-1.22 specificity precedence: most-specific wins regardless of
registration order, ambiguous overlaps panic at registration, use/mount stay
declaration-ordered barriers). TestIAMKeysBeatsWildcard now passes as a
FRAMEWORK property — subsystem order-ints are no longer load-bearing for
route precedence (their init-ordering role remains; removal rides the
composition-root refactor). Also swaps the 6 direct gofiber test-file imports
to the fork (production code was already 100% behind zip).
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Both semantic subsystems (clients/code, clients/knowledge) embedded via their own
static CLOUD_AI_API_KEY HTTP client — a side-channel credential separate from the
chat path, and the root of vectors:0 (the key was dead + entitlement-gated). Add
Embed() to types.AIClient and route both embedders through deps.AI — the SAME
client chat runs through, which authenticates with the IAM client-credentials
(M2M) token when no static key is set: ONE org/project-aligned identity, one way,
no static key to rotate. Emits gen_ai OTel spans so embeddings are observable end
to end like chat.
This is the ai-module leg of the canonical path
ingress -> gateway -> commerce -> ai module; metering + no-exempt + billing-on land
next on this seam. Supersedes the CLOUD_EMBED_* dedicated-provider detour.
- types.AIClient: + Embed(ctx, model, inputs)
- clients/aihttp: httpAI.Embed (raw POST reusing the static/M2M authed transport)
- clients/disabled, clients/rpc: stub Embed (fail-closed / not-wired)
- clients/code, clients/knowledge: embed through deps.AI; drop the static key path
- knowledge test: exercise the deps.AI seam (inject a deterministic fake embedder)
Decomplect the observability estate to exactly ONE top-level concept.
## 5 -> 1 registration collapse
The plane was FIVE separately-registered subsystems (o11yscope 69,
o11y-runtime 71, o11y-event-ingest 68, o11y-otlp-ingest 72,
o11y-trace-inproc 73) whose names leaked five public concepts (five config
toggles + five /v1/<name>/health routes). The k8s-style ordering was an
internal impl detail. They collapse to ONE
`cloud.RegisterWithShutdown("o11y", 69, mountO11y, shutdownO11y, HealthOwner)`:
mountO11y performs the ordered sub-mounts in-process (mountEventIngest ->
mountScope -> mountRuntime -> mountIngest -> mountTraceSink), so the PUBLIC
concept is a single `o11y`. Behavior is preserved EXACTLY — every specific
/v1/o11y/* route still registers inside the one order-69 mount, hence BEFORE
the upstream hanzoai/o11y wildcard (order 70), so Fiber's in-order match still
gives the specific routes precedence. The upstream module co-owns the `o11y`
name at order 70 (the wildcard); HealthOwner on this entry keeps /v1/o11y/health
registered exactly once (never a duplicate).
## Flat, version-less public surface (one /v1/, no nested /api/vN)
The upstream SigNoz engine version is an internal impl detail resolved inside
the handlers, never leaked into a route:
- /v1/o11y/vm/{query,query_range} (was /v1/o11y/vm/api/v1/*) — VM proxy; the
upstream VM api/v1/* path stays INSIDE the handler (queryRaw). SuperAdmin
gate + {up,sum(up),count(up)} allowlist unchanged.
- /v1/o11y/{query,query_range} (new, query.go) — the flat builder query;
resolves to the v3 engine route SERVER-SIDE (the version-less alias would
float to v5, which 400s the console's v3 composite payload), delegating to
the same gated runtime handler the wildcard uses.
o11y stays EMBEDDED in-process (the everything-binary); OTLP ingest + trace
sink stay opt-in. Registers exactly one `o11y` in clients (grep-verified),
alongside analytics/evals/usage/audit.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Absorb the thin clients/commercesvc mount wrapper AND the fail-closed
clients/commerce_local.go stub into the in-tree commerce library, and expose
a real in-process commerce.Client — one package now owns the commerce library,
its /v1/commerce subsystem, and its cloud.CommerceClient, self-registered via
cloud's inversion hooks (the clients/kms pattern; cf. #241 gatewaysvc→gateway).
cloud never imports commerce, so there is no cloud⇄commerce import cycle.
commerce.Client (client.go): a real in-process types.CommerceClient. The former
stub failed CLOSED on every CheckEntitlement (the Phase-2 gap). It now resolves
org→active-subscription→plan-tier→license-features from commerce's OWN models +
the @hanzo/plans vocabulary:
- org → namespace via commerce's canonical org.Resolve (the same binding the
money path uses), so a read can never cross tenants;
- active, unexpired subscriptions via a plain Status filter (matches commerce's
own read idiom, so grant- and payment-created subs are both found);
- plan tier → flat license features via a new plan.LicenseEntitlement seam that
runs @hanzo/plans' toLicenseFeatures (the vocabulary stays in one place, not
re-implemented in Go);
- Active iff the plan's features carry "licensing.product:<id>".
MONEY-SAFETY: Active:true only for a real active sub whose real plan really
licenses the product. Every unresolvable piece (commerce not co-resident, org
unresolvable, subscription query error, plans vocabulary unavailable) returns an
ERROR → the entitlements gate treats it as "cannot verify ⇒ 503"; a clean
"no plan licenses this product" is Active:false (→ 402), never an error and never
a fabricated grant.
Decomplect: commerce.Mount takes a MountConfig VALUE (Brand/Env/DataDir/Domain)
+ a logger — the values it uses, not the whole cloud.Deps bag; mountFromDeps is
the one place Deps is narrowed and carries the PCI Payments/Vault warnings.
Network path preserved: pickCommerceClient still selects the ZAP-RPC/disabled
client when commerce is NOT enabled in-process (CLOUD_COMMERCE_ZAP_ADDR). This
fold does NOT force the live in-process cutover — that stays operator-gated.
Deleted: clients/commercesvc, clients/commerce_local.go(+test), the legacy
//go:build cloud clients/commerce/mount.go(+ its two tests), and the redundant
cmd/commerce --cloud boot path (cloud_boot.go/cloud_stub.go) — one way to serve
commerce inside cloud: the folded subsystem.
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Concurrent RO opener reads all records while the serialized writer appends;
writer chain stays intact and a fresh RW re-open recovers the head with no
fork/gap. Audit-store analog of clients/kms TestReaderReadOnlyRoundTrip.
Locks in the reader/writer shareability invariant the cloud HA carve depends
on. Test-only; zero runtime change (writer path byte-identical).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The code + knowledge embedders read CLOUD_AI_BASE_URL/CLOUD_AI_API_KEY,
which are ALSO the chat/synth AIClient's config (config.go:308). That
braided embeddings onto the IAM-gated api.hanzo.ai path, whose /embeddings
403s customer keys and 401s gateway keys — only a superadmin JWT passes —
so the vector tier indexed vectors:0 platform-wide.
Split embeddings onto dedicated CLOUD_EMBED_BASE_URL / CLOUD_EMBED_API_KEY /
CLOUD_EMBED_MODEL, each falling back to the shared CLOUD_AI_* for back-compat.
Both semantic subsystems (clients/code + clients/knowledge) read the same
three vars — one embed config, orthogonal to chat. Deployment points them at
DigitalOcean GenAI (bge-m3 @1024-dim, first-party, not IAM-gated) via the
DO_AI_API_KEY already in cloud-api-llm-keys; chat stays on CLOUD_AI_*.
Secrets are write-only: toAppView masks a secret value to "" on read, so a client
editing env re-submits kept secrets with an empty value. sealSecretEnv now treats an
empty secret value as "keep the already-sealed value" — it is preserved, not resealed
to empty (which wiped the KMS secret). Only a non-empty value seals (and requires KMS).
Makes the masked-read -> setEnv round-trip a no-op for untouched secrets instead of a
data-loss footgun; unblocks the console env editor's Keep path. Adds a focused test.
The gateway subsystem package carried an "svc" suffix as a naming habit
only: there is no clients/gateway to collide with, and the package imports
no "gateway" library, so it renames cleanly to the plain domain name.
- git mv clients/gatewaysvc → clients/gateway (history preserved)
- gatewaysvc.go / gatewaysvc_test.go → gateway.go / gateway_test.go
- package gatewaysvc → package gateway; Mount() error strings + doc comment
updated to gateway.*
- update the blank import in subsystems/subsystems.go and the doc comments
in deps.go, middleware_edge.go, clients/gatewaypolicy/policy.go
commercesvc is intentionally NOT renamed in this PR. clients/commerce
already exists in-tree — the embedded upstream Hanzo Commerce library
absorbed by #114 (package commerce, ~1583 files) — and commercesvc.go
imports it as github.com/hanzoai/cloud/clients/commerce. Moving
clients/commercesvc → clients/commerce is therefore a directory-level
collision, not an import-alias shadow; re-aliasing does not free the path.
Renaming commerce cleanly needs the operator to decide the target name (or
relocate the embedded library), so it is flagged for follow-up, not forced.
Build: GOWORK=off CGO_ENABLED=0 go build ./... → exit 0
Test: GOWORK=off CGO_ENABLED=0 go test ./clients/gateway/... \
./clients/gatewaypolicy/... . → ok
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The console's /metrics and /status SuperAdmin infra-health board read
VictoriaMetrics `up{}` via the console's Next.js `/telemetry/[...path]` route.
That server route is STRIPPED by the static-export embed (cloud go:embeds
console as output:'export'), so the browser call 404s — the last console error.
Add a same-origin, SuperAdmin-gated VM read proxy on the cloud `/v1` API that
serves the exact three queries the board issues, registered at o11yscope order
69 (before the o11y wildcard at 70):
GET /v1/o11y/vm/api/v1/query?query=up
GET /v1/o11y/vm/api/v1/query_range?query=sum(up)|count(up)&start&end&step
Security follows scope.go's contract ("the client never supplies a raw query"):
- admin(c) gate (X-User-IsAdmin, reserved admin org) — 403 for everyone else.
- The ?query param is ALLOWLISTED to exactly {up, sum(up), count(up)}; anything
else is 400. Range args (start/end/step) validated as positive integers. No
generic PromQL passthrough — this can never become an exfiltration/DoS surface.
- VM's native Prometheus envelope is returned VERBATIM (c.Bytes) so the console's
parseInstant/parseRange work unchanged. Reuses the existing newVMClient()
(O11Y_VM_URL); adds vmClient.queryRaw for the verbatim forward.
Pairs with console 1beb6fca9 (telemetry.ts repoint). Verified: go build + go vet
clean on clients/o11y.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Every co-resident subsystem that spoke the commerce billing S2S surface over HTTP
to the standalone pod (commerce.hanzo.svc:8001) now dispatches straight into the
in-tree commerce handler (#114) — no socket, no serialization change. This is what
lets the standalone be retired.
- New seam clients/commerceinproc: commercesvc.Mount publishes the embedded
commerce http.Handler (now carrying /v1/billing via api.Route) once at boot; a
self-routing http.RoundTripper dispatches each S2S request to it in-process when
co-resident, else falls back to plain HTTP (the pre-#111 split-deploy behavior).
Behaviour-preserving: same path, same Bearer+X-Org-Id headers, same body/status
bytes either way — proven by commerceinproc_test (in-process dispatch, HTTP
fallback, base resolution).
- Converted: clients/{billing,account,admin,referrals,authors,affiliates,usage} —
each keeps its OWN request-building + tenant subject-pinning (the security
boundary is untouched); only the transport swaps. account keeps its shared
httpClient for HUSD EVM JSON-RPC and routes ONLY commerce calls in-process.
- The request-edge METERING gate (build.go buildMeteringClient) debits the
in-process handler; base pinned non-empty when co-resident so the gate never
silently no-ops (free-money hole) even after CLOUD_COMMERCE_HTTP_URL is dropped.
GATES: go build ./... + cmd/cloud (sqlite tags) green; commerceinproc + all 8
converted packages' tests pass. LIVE money-parity gate (in-process vs standalone
balance/deposit/usage on shared Postgres + metering debits) precedes retiring the
standalone — that live cutover is gated, not in this commit.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Move hanzoai/commerce out of the external-dep seam and INTO the cloud module,
so /v1/commerce (and, next, /v1/billing) is served by one repo, one binary, one
way. No external github.com/hanzoai/commerce* require remains.
Absorbed the EXACT versions cloud already compiled (behavior-preserving — not
main), copied from the module cache so the in-tree tree == what the live binary
built:
- github.com/hanzoai/commerce v1.46.40 -> clients/commerce
- github.com/hanzoai/commerce/metering v0.1.4 -> clients/commerce/metering
- github.com/hanzoai/commerce/thirdparty/ethereum v1.40.0 -> clients/commerce/thirdparty/ethereum
- ONE Go module: removed the three nested go.mod/go.sum; commerce's deps merged
into cloud/go.mod via `go mod tidy` (MVS keeps the highest existing patch — the
single major conflict, luxfi/zap, already resolved to cloud's v1.2.1 in the old
build, so nothing recompiles differently). New direct deps discovered by tidy:
huandu/facebook, square-go-sdk/v3.
- Imports rewritten repo-wide: github.com/hanzoai/commerce* ->
github.com/hanzoai/cloud/clients/commerce* (1103 files: the absorbed tree +
cloud's own build.go/deps.go/middleware_billing.go/clients/{bots,functions}/…
metering importers). commercesvc leaf now imports the in-tree package.
- Excluded the app/ frontend pnpm workspace (59M, NOT referenced by any .go —
the built ui/dist + billing/ui/dist + checkout/ui/dist that Go //go:embed's are
kept). //go:build cloud files kept (excluded from the plain build as before).
GATES (all green):
go build ./... -> ok
go build -tags "libsqlite3 sqlite_fts5" ./cmd/cloud -> ok (633MB binary boots)
go test ./clients/commercesvc/... ./clients/commerce/{,datastore,api/billing} -> ok
go test -run 'Billing|Metering|SpendCap|Commerce' . -> ok
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
cloud pinned beego at a PSEUDO-VERSION
(v2.4.1-0.20260710093857-0ad99bdf8b90). A pseudo-version forces `go` to
resolve the dep via a live VCS fetch-by-commit-hash into the shared ARC
GOMODCACHE cache/vcs. Parallel containment jobs racing that shallow bare
repo hit `fatal: shallow file has changed since we read it` → intermittent
FAIL, which has forced admin/green-locally squash-merges (#235, #236) and
defeated the containment gate.
beego HEAD (0ad99bdf) is exactly one benign commit past tag v2.4.0 (silences
a missing-default-app.conf stderr warning, beego#29). Cut that commit as a
proper patch tag v2.4.1 and pin to it — byte-identical code, but `go` now
fetches the immutable refs/tags/v2.4.1 ref instead of shallow-fetching a
bare commit, so no cache/vcs race. beego is private (no module-proxy), but
the tag ref path is deterministic and does not trip the shallow-file check.
go.sum regenerated via `go mod tidy` (authentic v2.4.1 hash). `go build ./...`
green. iam stays on tag v2.3.10 (already a tag, not a pseudo-version).
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
cloud installs the ONE global tracer provider wired to the o11y in-process trace
sink, but in in-process-sink mode it sets no OTLP/ZAP exporter endpoint. The
embedded hanzoai/ai module then found no endpoint and DISABLED its gen_ai span
emit, so the gen_ai plane in o11y_traces was dark (cloud's own request spans
flowed; every LLM call's gen_ai span was gated off).
After otel.SetTracerProvider, call aiobject.AdoptHostTracerProvider() so ai emits
every gen_ai span through THIS provider -> the o11y in-process sink -> o11y_traces.
Repins ai to the branch carrying AdoptHostTracerProvider. The OTLP-env unset stays
as orthogonal defense against any other lib forking a competing provider.
ai: github.com/hanzoai/ai v1.804.1 -> v1.804.2-0.20260710172916-723e03f81bca
Simple subsystems use one-line cloud.Mount[S]. Subsystems whose Mount is more
than build+routes (background reconciler, package-global for cross-package hooks,
shutdown cancel — e.g. platform) construct the Service value directly with
cloud.NewBase + &cloud.Service[state]{...} and wire routes via Handle — still the
ONE generic type, one Base derivation.
The ONE server abstraction: a subsystem is cloud.Base (shared deps derived once:
Log/KMS/Bill/Brand/Env/Domain/DataDir) + its own typed State. Handlers are FREE
FUNCTIONS func(*Service[S], *zip.Ctx) error bound with cloud.Handle — so a package
declares NO service/receiver type, only its State (plain data) and its handlers.
cloud.Mount[S] is the one generic entrypoint (build state, wire routes).
Pike minimalism: generics carry the state type, functions carry behaviour. Kills
the 'type svc struct{ …re-plumbed deps… }' shape copied ~40×.
Proven on clients/usage: svc type gone, log via embedded Base (s.Log), commerce
reader in State (s.State.commerce), all handlers/helpers free functions. Build +
vet + test green, routes + tenant isolation unchanged.
* feat(#105): un-stage commerce — serve /v1/commerce in-process from the one binary
Phase 2 final step (tasks #96 → #105). commerce was the last staged
in-process subsystem; drop it from stagedSubsystems so the mount-all default
serves /v1/commerce (+ /_/commerce) from the cloud binary instead of proxying
to the standalone commerce pod. iam + ingress STAY staged (the IAM embed
corrupts its own Beego bootstrap under mount-all).
commerce owns the money path, so the cutover is data-neutral BY CONSTRUCTION:
the authoritative stores stay put and shared — balances/deposits/credits live
in Hanzo SQL (SQL_URL), analytics in DATASTORE_URL, blobs in S3 — and the
in-process commerce.Embed reads the SAME stores via the SAME backend env the
standalone CR carried. Only the small per-org merchant SQLite + tenant `base`
tree migrate into the cloud data dir. No money is copied or split.
Tests: go test -run TestEnabled ./ → ok (staged contract: commerce now mounts
under the empty-Enable default; iam/ingress still gated to explicit CLOUD_ENABLE).
Also gofmt: fixes a pre-existing struct-literal misalignment in LoadConfig.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
* fix(commercesvc): isolate in-process commerce data under {DataDir}/commerce
In-process commerce writes orgs/ and base/ under its DataDir. cloud sets
deps.DataDir=/var/lib/cloud, where cloud ALSO owns /var/lib/cloud/orgs (its
own per-org subsystem SQLite, HIP-0302) and /var/lib/cloud/base (its Base/IAM
store) — verified live. Sharing the root would open two apps on the same
SQLite files (base/data.db) and corrupt them. Always nest commerce under
{deps.DataDir}/commerce (was only the empty-DataDir fallback), keeping the
commerce ledgers physically separate on the same cloud-api-data PVC. This is
the target path the #105 data migration copies the per-org merchant SQLite +
tenant base into.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
---------
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The consumer half of the SBOM datastore lane. The registry is the source of
truth: CI produces the CycloneDX SBOM and `cosign attach`es it to the image
digest. Cloud now PULLS that attached artifact and materializes its components
into the global hanzo.sbom_component table — no CI push-ingest.
clients/sbom/pull.go (new): registry pull with go-containerregistry (pure-Go,
no binary deps). Resolves ref→digest, locates the CycloneDX SBOM by the cosign
SBOM tag (sha256-<hex>.sbom) first, OCI 1.1 referrers second, and parses it
through the SAME parseComponents the POST /v1/sbom path uses (one flattener).
Triggers:
- Pull-on-miss: GET /v1/sbom/{ref} with 0 rows and an image ref pulls the
attached SBOM, upserts it, and rereads FINAL — the console goes live with no
console change (the panel already GETs by repository:tag).
- Deploy-time: platform applyLive fires sbom.Prefetch(ref) async when a
deployment goes live (best-effort, idempotent; platform→sbom, one direction).
go.mod: + github.com/google/go-containerregistry v0.21.7 (direct), authentic
go.sum via go mod tidy.
Tests: hermetic end-to-end over a real in-memory OCI registry (cosign tag +
OCI referrers + no-attachment + bare-digest), all green with -race. Plus an
env-gated live end-to-end (pull_live_test.go) proven against registry:2 + real
ClickHouse serving a real cyclonedx-gomod SBOM: GET → pull-on-miss → production
pullSBOM → parse → INSERT → reread → 200 with the real components + cache hit.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
A push landed on the embedded git server (clients/git) now fires a build
for every app that tracks that repo+branch — no GitHub, no Actions.
- build.go: GitPushEvent + RegisterPushBuilder/OnGitPush — the same
package-level inversion as kmsClientFactory, so git never imports
platform and there is no git<->platform cycle.
- clients/git/smart_http.go: after a push lands (post-metering),
firePushBuilds fires OnGitPush for each branch ref that advanced
(tags + deletes skipped). Best-effort: a trigger failure is logged,
never fails the push the client already committed.
- clients/platform/push.go: buildFromPush resolves every git-source app
whose RepoURL+branch matches the pushed ref and launches a build via
the ONE shared build-launch core.
- clients/platform/deploy.go: extract startGitBuild (ctx-only) as that
single core; deployGit maps it onto the HTTP deploy, buildFromPush onto
the push trigger — one build-launch path, no duplication.
- clients/platform/validate.go: the cloud's own embedded-git apex
(deps.Domain) is always a trusted build source, so a self-hosted-git
app builds with no env — host all repos on our own git, not GitHub.
Tests: TestPushFiresBuildTrigger (real go-git push -> OnGitPush fires with
org/repo/branch/commit/cloneURL), TestBuildFromPush_{LaunchesMatchingApp,
NoMatchIsNoop,IgnoresImageApp}. Full platform + git suites green.
Jul-5 team-go→cloud migrated workspaces threw "Confirmed social identity is attached to the wrong person" on transactor connect: the confirmed hanzo:<account> SocialIdentity was still attached to a team-go-era Person id, not the deterministic person-<account>. reconcile() now runs remapMigratedSocialIds(uid) to re-point it — migrated-only, idempotent, non-destructive. CI red check (controlplane-containment) is an unrelated shared-runner Go module-cache flake on hanzoai/beego; fix is clients/team-only and green locally (go build ./..., make build, go test/-race/vet/gofmt).
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Set O11Y_AUTHZ_PROVIDER=local so the in-process o11y authorizes org-scoped,
gateway-authenticated users locally instead of round-tripping to an external IAM
Casbin enforcer the one-binary has no credentials for. That round-trip was 401ing
every /v1/o11y read (provision Grant -> add-policy authz_unavailable), breaking the
console overview-metrics widgets on ~9 secondary product pages. Same enforced
policy; tuples in-process. Bumps o11y 1.5.9 -> 1.5.10.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
v1.5.8 allows digits in TypeRole selectors, unbreaking the built-in o11y-admin role
grant that panicked (→500) on EVERY authenticated embedded o11y data read. Completes
the o11y-telemetry fix (v1.5.6 aliases + v1.5.7 /api passthrough + v1.5.8 grant):
/v1/o11y/{query_range,services,rules,dashboards} now resolve for the console
overview-metrics widgets.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
v1.5.7 routes /api/*-prefixed paths straight to the router in the ExternalPath wrapper,
fixing the double-strip that 404'd EVERY embedded o11y data call. With v1.5.6's
version-less aliases, /v1/o11y/{query_range,services,rules,dashboards}+/metrics now
resolve — fixes the console overview-metrics widgets on studio/gateway/cli/registry/
desktop/console/dashboards/alerts/metrics.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
v1.5.6 registers the version-less /api/<resource> aliases on the app.Server router
the embedded runtime uses (o11y e01015954), so the console's /v1/o11y/{query_range,
services,rules,dashboards} + /metrics resolve instead of 404 — fixes the o11y-telemetry
overview widgets on studio/gateway/cli/registry/desktop/console/dashboards/alerts/metrics.
Also folds in the v1.5.5 C1 cross-tenant llmobs read fix.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The prior fix degraded the orgRouters-error path, but zt's gate() fail-closes with
503 BEFORE that when ZT_CLIENT_ID/SECRET are unset — so /networks + /edge still 503'd
a console error on every load. Short-circuit the two READ handlers to an honest-EMPTY
list (200) when unconfigured, ahead of the gate. WRITES keep the fail-closed 503.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Every `hanzo` control command (version, whoami, config, apps, login, …) printed
three lines of server-mode init chatter first:
init global config instance failed ... open conf/app.conf: no such file...
failed to load persistent registry signing key: KMS_SERVICE_TOKEN ... required
generating ephemeral registry signing key for non-production runtime
These come from dependency package init() functions that run before main(), so
main.go's client-vs-server branch cannot gate them, and Go's alphabetical
package-init order (beego < cloud < iam) means no cloud-side os.Stderr swap can
pre-empt them — verified with an init-order probe. Masking the output would also
leave a wasteful KMS fetch + RSA keygen firing on every CLI call. Fixed at the
source instead:
- hanzoai/iam#117: registry signing key resolved lazily (sync.Once) at its use
sites, not in package init(); server token paths keep identical semantics,
the CLI never touches KMS or generates a key.
- hanzoai/beego#29: the benign conf/app.conf probe is silent when the default
file is absent (the CLI case), loud only on a present-but-broken file.
Bump both deps to the fixed versions and document the quiet-output invariant at
the CLIENT MODE gate in cmd/hanzo/main.go.
Result: `hanzo version|whoami|config path|apps --help` emit ZERO stderr; server
subcommands (`hanzo ai --help`, serves) initialize unchanged.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
All four are folded in-process and validated on cloud-unified-canary
(v1.786.167): commerce embedded, captable/sign/dataroom "mounted in-process
(goja + per-tenant Base)", /v1/{commerce,captable,sign,dataroom}/health all 200.
Dropping them from stagedSubsystems makes the mount-all default serve them, so
the main cloud (empty CLOUD_ENABLE) runs them natively and their standalone
Postgres/Next pods retire. iam + ingress stay staged (IAM embed corrupts its own
bootstrap under mount-all; iam served by the standalone pod).
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
console.hanzo.ai /indexer, /oracles, /networks, /edge fired 502/503 console errors
because these list handlers returned the upstream error when the chain indexer /
price-feed oracle / ZT controller is unreachable or unconfigured (ZT_CLIENT_ID
optional). Same graceful fold already shipped for visor clusters/machines/gpus:
log + return an honest-EMPTY list (200). An empty list is honest (you have none),
never fabricated; ZT WRITES stay fail-closed. Clears the console errors on 4
secondary product pages for every org without those backends deployed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
"console" is just our cloud FE name; there must be NO /v1/console/* API domain.
Rename clients/console → clients/account and re-home every route onto its REAL
domain, forwards-only (no /v1/console aliases, no compat shim):
keys GET/POST/DELETE /v1/console/keys → /v1/iam/keys
onboard POST /v1/console/onboard → /v1/iam/onboard
csrf GET /v1/console/csrf → /v1/csrf
embed-status GET /v1/console/embed-status → /v1/embed-status
topup POST /v1/console/topup/wallet → /v1/commerce/topup/wallet
health /v1/console/health → dropped (generic /v1/<name>/health)
waitlist /v1/console/waitlist → dropped (SPA uses clients/base /v1/waitlist)
billing/commerce bridges /v1/billing/*,/v1/commerce/* paths unchanged, relocated
The same handlers, CSRF protection (requireCSRF), per-IP rate limiting, and the
VALIDATED-principal tenancy are preserved — only the PATHS change.
IAM wildcard ordering: keys/onboard live on /v1/iam/*, which clients/iam (order 50)
mounts as a WILDCARD. The package registers TWO subsystems so the specific routes win
Fiber's first-match scan: `account` (order 48) mounts /v1/iam/{keys,onboard} + /v1/csrf
+ /v1/embed-status + /v1/commerce/topup/wallet BEFORE the wildcard (and before the
commerce embed at 100); `account-bridge` (order 122) mounts the /v1/billing/* +
/v1/commerce/* catch-all bridges AFTER clients/billing (121) + the commerce embed (100).
Both share one svc + a process-wide CSRF key so a /v1/csrf token verifies on the bridge
writes. Proven by TestIAMKeysBeatsWildcard (native /v1/iam/keys beats the wildcard;
/v1/iam/oauth/token still falls through) and TestRegisteredOrders (account<50, bridge=122).
waitlist.go/waitlist_test.go deleted; httpClient relocated to topup.go. All 54 tests
pass; go build ./... green.
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Fold hanzoai/dataroom (Papermark fork: Next.js + Prisma + Postgres) FULLY into
the unified cloud binary via the gojahost pattern (HIP-0106, task #101 / epic
Postgres, no Next.js.
REUSE the shared binding, don't build a second. clients/dataroom runs the
self-contained dataroom goja bundle (byte-identical to hanzoai/dataroom/goja/
bundle.js, go:embed) on the REUSABLE clients/gojabase host — the SAME RW-Base
binding captable (#97) pilots and esign (#100) reuses — which injects
__db/__newId/__now, opens one SQLite file per tenant, and runs each dispatch in
one transaction (commits iff status<400). The leaf adds only: the per-tenant
Schema (schema.go), the object-storage seam for document bytes, a bcrypt HostFn
for link passwords, and the public link->org index. Zero domain logic in Go.
gojabase gains ONE generic, domain-free seam — Config.HostFns — the extension
point esign/dataroom both reach for (dataroom injects __bcrypt; the reserved
__db/__newId/__now always win). captable is unaffected (nil HostFns).
Storage: document BYTES go through the cloud object-storage seam (deps.VFS — the
S3/SeaweedFS data plane, a Go storage host-fn in the leaf, NOT local FS); the
bundle persists only the opaque key. View-analytics events (page-by-page
tracking) are Base rows in the tenant DB.
Auth: admin routes require a validated cloud principal (principal.Tenant -> org);
public viewer routes resolve their org from the link index. Passwords hashed with
bcrypt in Go — never plaintext. Registered order 134, HealthOwner, STAGED behind
CLOUD_ENABLE. go.mod unchanged.
Proof (clients/dataroom/flow_test.go, in-process over real per-tenant Base + the
VFS seam): create dataroom -> upload document (bytes to storage) -> attach ->
create email+password-gated share link -> open as public viewer -> authenticate
(wrong password 401, disallowed email 403) -> record per-page views -> analytics:
{"pages":[{"pageNumber":1,"views":2,"totalDuration":6000,"avgDuration":3000},
{"pageNumber":2,"views":1,"totalDuration":900,"avgDuration":900}],
"totalPageViews":3,"totalViews":1}
plus viewer byte download round-trip and cross-org isolation. Live binary boots
with CLOUD_ENABLE=dataroom, mounts in-process, health 200, admin 403 fail-closed.
Companion bundle PR: hanzoai/dataroom#6.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Fold the esign product (Documenso fork — open-source DocuSign) FULLY into the
unified cloud binary, per HIP-0106 (epic #96, task #100). Cloud serves /v1/sign/*
itself — the TS domain on dop251/goja backed by per-tenant Hanzo Base/SQLite. No
Next.js, no Prisma, no Postgres. Reuses the SAME clients/gojabase RW-Base binding
the captable pilot (#96) established — ONE binding, not a second.
clients/gojabase: add an additive, backward-compatible Config.HostFns passthrough
— extra native host globals injected per dispatch alongside __db/__newId/__now.
esign uses it for __pdf; captable is unchanged (tests green).
clients/sign (leaf on gojabase):
- schema.go — the per-tenant sign.db DDL (gojabase Config.Schema).
- signer.go — THE HARD PART as Go host-functions injected via HostFns as
__pdf = { stamp (pdfcpu renders field values onto the PDF), sign (real
x509/PKCS#7 seal via digitorus/pdfsign) }; signer sourced from KMS PEM,
persisted PEM, or a self-signed dev cert. Signing orchestration stays TS.
- sign.go — Mount + route table; owner routes gated by principal.Tenant,
recipient token routes org-in-path. Registered + STAGED behind CLOUD_ENABLE
(config.stagedSubsystems); retires the standalone esign pod on cutover.
- sign_test.go — end-to-end wire proof: create→recipient→fields→send→sign→
complete seals a REAL signed PDF (/ByteRange,/Type /Sig,PKCS7) + full audit,
per-tenant Base-backed, with cross-tenant isolation.
Bundle: github.com/hanzoai/sign (goja/bundle.js) — the ESM-free domain port on
the gojabase contract (__db/__newId/__now/__pdf, handle{route,params,query,orgId,
body}, one txn per dispatch). go.mod additive; pdfcpu pinned v0.11.0,
hhrutter/pkcs7 v0.2.0 (no bumps).
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Ship the crown-jewel security fix: o11y v1.5.5 (org-scoped llmobs span views — closes
the cross-tenant read exposed by the Observe->/v1/o11y repoint) + ai v1.804.1 (BYO
billed_cost). The leaky observation read is already deleted upstream (#217); this repin
ships the org-scope via the cloud-embedded o11y runtime.
(Coordinator asked for v1.786.162 but that tag + through v1.786.165 were already cut on
origin; this is the next free tag.)
Crown-jewel security repin off the pseudo-versions to the merged release tags:
- o11y v1.5.4 -> v1.5.5: the llmobs span-view SQL is now org-scoped
(gen_ai.hanzo.org_id = <validated tenant>, fail-closed) — closes the cross-tenant
read the Observe->/v1/o11y repoint exposed. Cloud EMBEDS the o11y runtime
(clients/o11y), so this is the ship vehicle for the fix.
- ai <pseudo> -> v1.804.1: gen_ai span emits _o11y.gen_ai.billed_cost alongside
total_cost (BYO invoice reconciliation) + the enriched gen_ai span.
The leaky cloud_usage-as-observations read stays DELETED (upstream via #217); the
metering warehouse (cloud_usage via /v1/usage + the #218 /v1/evals/metrics dashboard)
is untouched. No go mod tidy (luxfi/keys go.sum fragility) — only the two repins.
Build: go build ./... green. Test: clients/eval + clients/o11y green.
Cloud now serves /v1/captable/* ITSELF, per tenant, on Base/SQLite — the PILOT of
epic #96 (fold the Captable,Inc app into the unified binary; drop Next.js/Prisma/
Postgres). Where clients/plan + clients/pricing host a read-ONLY @hanzo catalog in
goja, captable hosts the tRPC business LOGIC (ported to a self-contained goja
bundle in github.com/hanzoai/captable) and gives it PERSISTENCE over per-tenant
Base/SQLite. The bundle carries logic; the Go host carries storage.
REUSABLE Base-goja binding (clients/gojabase) — the deliverable esign (#100) +
dataroom (#101) rebase onto. It is the storage-bearing sibling of clients/goja
(the pure JS engine): given a Bundle + a per-tenant Schema (DDL) + DataDir, it
- opens ONE SQLite file per tenant (lazy, migrated once, cached; slug-contained),
- injects per dispatch a tenant-bound __db bridge (query/exec) + __newId + __now,
- runs globalThis.handle inside ONE transaction that commits iff status<400 and
handle didn't throw (atomic multi-statement mutations for free), and
- carries ZERO domain logic. clients/goja gains DispatchWith (per-call native
globals) as the read-WRITE extension of the read-only plan/pricing path.
clients/captable leaf: go:embed'd bundle (hanzoai/captable.Bundle) + the per-
tenant schema (Prisma model → SQLite DDL) + a company seed (OnOpen) + the
/v1/captable/* zip routes. Org resolves from the VALIDATED principal
(principal.Tenant), never a client header; that org selects the DB file AND
scopes every row. Registered order 133; STAGED behind CLOUD_ENABLE (joins iam/
ingress/commerce in config.stagedSubsystems) so main stays shippable and the
standalone captable service keeps authority until the phase-2 cutover.
Full fold over Base: stakeholders, share classes, equity plans, securities
issuance (shares + options), share transfers (full + partial, atomic), SAFEs +
convertible notes, rounds + investments (a priced round issues shares and
dilutes), and a computed cap table (fully-diluted ownership, per-class
authorized-vs-issued, convertibles + rounds summary).
Proven:
- clients/gojabase: RW round-trip, per-request rollback (caught-500 + raw-throw),
per-tenant isolation, OnOpen seed, slug traversal-containment (real SQLite).
- clients/captable: the REAL embedded bundle vs REAL SQLite through the whole
lifecycle (create stakeholder → issue share class → issue shares → priced round
+ investment dilution → transfer → cap table), and an HTTP wire test via the
zip/Fiber test client (trusted headers) proving create→read-back, issuance→cap
table, the bundle's OWN 404 (not a proxy 502), the 403 principal gate, and
cross-tenant isolation.
- Live binary boots under CLOUD_ENABLE=captable and serves /v1/captable/health
200 in-process (a proxy would 502).
go build ./... + go build -tags "libsqlite3 sqlite_fts5" ./cmd/cloud green.
go.mod pins github.com/hanzoai/captable at its merged main commit; go.sum authentic.
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The server side of the ex-/v1/arcd surface: one native build API on the
runner fabric that hanzo build, git-push-to-deploy, and cloud's own
self-release all call — no GitHub builders, no Actions.
- runner.go: POST /v1/runner. Privileged (caller supplies the output
image), so gated by a constant-time build-callback token AND an
image-ref allowlist (ghcr.io/{hanzoai,luxfi,zooai}/*). Fails closed
when no token is configured.
- k8s.go: launchDirectBuild — validated at the same choke point as the
tenant build; plus buildFrontendCmd + buildJobSpec extracted so tenant
and direct builds share ONE Job spec and ONE frontend selector.
- hanzoai/pack is the default BuildKit frontend (zero-config, gateway.v0);
a Dockerfile is the explicit escape hatch (dockerfile.v0).
Tests: token 503/403, missing-field 400, image-allowlist 403, happy-path
202. Full platform suite green; whole cloud module builds.
Before: subsystems opened SQLite inconsistently. clients/code used the good
per-ORG-file pattern ({DataDir}/orgs/{slug}/code.db, resolved per-request), while
git/functions/tracker each opened ONE shared DB at Mount ({DataDir}/git.db,
functions.db, tracker.db) scoped only by an org column per-row — a decomplected
tenancy gap where the physical boundary was a single file for every tenant.
After: ONE resolver — cloud.TenantDB(dataDir, org, project, subsystem) — is the
single way any subsystem opens a tenant SQLite DB. Path convention:
project-scoped: {DataDir}/orgs/{orgSlug}/projects/{projectSlug}/{subsystem}.db
org-scoped: {DataDir}/orgs/{orgSlug}/{subsystem}.db
It MkdirAll 0700s, opens via the sole "sqlite" driver (github.com/hanzoai/sqlite),
applies the shared single-writer + WAL pragmas, and folds org/project through the
injective SanitizeOrg slugger so distinct tenants can never share a file and no
segment can traverse. A generic cloud.TenantStore[T] caches per-tenant stores
(opened once each) so the hand-rolled per-subsystem map is DRY'd into one value.
SanitizeOrg (the one injective org-slug normalizer) moves to the root cloud
package beside OrgHasUnsafeRune and TenantDB; provisioning.SanitizeOrg delegates
to it, byte-identical, so S3/KMS/knowledge slugs are unchanged.
Subsystems migrated onto the resolver:
- code: adopts the helper; stays ORG-scoped (no project axis) — same path.
- git: single-shared git.db -> per-ORG file. Kept org-scoped (NOT project)
because /v1/git/usage is a deliberate org-wide rollup across every
project; the project stays a row column.
- functions: single-shared functions.db -> per-ORG file (no project axis).
- tracker: single-shared tracker.db -> per-(org, IAM-project) file — tracker is
project-scoped (principal.Project); its KEY-based projects are rows
WITHIN each per-project file.
- tasks: left as-is — it opens NO SQLite (delegates to the shared durable
engine owned by durable.go), so there is nothing to migrate.
Data safety: the single-shared -> per-tenant switch is fail-closed (an invalid
org/project errors rather than falling through to another tenant's file). These
are new subsystems with little/no production data; existing rows in a prior
shared *.db would live under {DataDir}/{subsystem}.db and are NOT auto-migrated —
a deployment carrying such data must relocate rows into the per-tenant files
before cutover. No silent data drop.
Tests prove isolation: two orgs -> two files, no cross-read; two projects under
one org -> two nested files; project-scoped path nests; the cache opens each
tenant once. Root tenantdb_test.go plus per-subsystem end-to-end file-isolation
tests (git/functions org, tracker project).
Co-authored-by: zeekay <ai@hanzo.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
CONSOLE_CACHEBUST was the cloud sha, so a console-only push could not trigger a
fresh embed without a cloud commit — the whole point (freshness) leaked between
cloud pushes. Resolve hanzoai/console main HEAD (git ls-remote, extraheader
cleared since actions/checkout's GITHUB_TOKEN 404s cross-repo; gh absent on the
runner) and use it as the cachebust; fall back to the cloud sha if empty. A
console change now moves the cache key on its own.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* controlplane inc-2 seam (a): per-pod ML-DSA-65 identity keys
Replace the symmetric-HMAC proof-of-possession with per-pod asymmetric ML-DSA-65 (github.com/luxfi/crypto/mldsa, MLDSA65, FIPS 204 Level 3) identity keys — the same key material that becomes the cert-signing key in seam (c). idKey is now a fresh random keypair (crypto/rand), NEVER seed-derived; the registry stores the public key; signPoP uses the FIPS 204 5.2 deterministic variant with a domain-separation context; VerifyPoP is asymmetric VerifySignatureCtx.
Acceptance: the two byzantine-safety tests TestRed_B_DerivablePoPForgesHonestLeg and TestRed_B_LoneNodeForgesFullQuorum are un-skipped and now GREEN (an attacker rebuilding a pod's custody from public inputs gets a different key whose PoP fails); TestSafety_RogueAndForgedLegs_Rejected is re-armed to forge a CORRECTLY-DERIVED leg under a real attacker key (not a bit-flip) and still rejects it.
Scope: seam (a) only. ProductionBCCSigningReady() stays false and all other stubs remain — the flag flips later in the same change that lands the real cert (seam c) and deletes the stub types. Full suite green (zero skips), -race clean, gofmt clean, containment intact (no-tag build still matches no packages).
* controlplane seam (a): R3 — doc.go PoP narrative now reflects discharged crypto
RED review R3: refresh the stale CLASS-B caveat + ShareCustody stub-catalog entry so prose matches the code. Per-pod identity keys are real random ML-DSA-65 (not seed-derived); the two TestRed_B_* are un-skipped + green; only the z-share / cert crypto remains stub (ProductionBCCSigningReady stays false until seam c). Doc-only; no code change.
Adds the write path for LLM-observability events (traces/observations/scores)
that the retired console-worker (Node BullMQ->Valkey->Datastore) used to
provide, folded into the cloud binary as a normal o11y subsystem.
- POST /v1/o11y/ingestion (validated tenant via principal.Tenant) -> parse batch
-> route by type -> per-table batch insert into the Hanzo Datastore via the
branded github.com/hanzoai/datastore-go/v2 client (promoted to a direct dep;
first direct user in cloud). ZERO clickhouse imports/identifiers -- datastore-go
brings the same upstream ch-go line (MVS-unified v0.71.0) the SigNoz o11y query
runtime uses, so it coexists in one binary (verified).
- Grounding: the embedded o11y runtime (github.com/hanzoai/o11y) is a SigNoz fork
serving INFRA o11y and mounts app.All("/v1/o11y/*") at order 70; it has no
LLM-obs ingestion. This is a cloud-native SPECIFIC route registered at order 68
so Fiber's in-order match binds it AHEAD of the wildcard (same rule scope.go
uses for /v1/o11y/{logs,metrics,status} at 69). A route at the OTLP-ingest
order (72) would be swallowed by the proxy.
- Oversized event bodies overflow to object storage (blob ref stored inline),
threshold via CLOUD_O11Y_INGEST_BLOB_BYTES (default 1 MiB).
- Always mounted (no feature flag); the Datastore (O11Y_DATASTORE_DSN) is a
required dependency. Fail-soft: no DSN / unreachable datastore -> unmounted,
never blocks boot. Inert in prod until the console producer repoints here.
Tests: go test -tags 'cloud cloud_mount' ./clients/o11y/... green (8 cases:
routing, batching, blob overflow/disabled, sink+blob error propagation, threshold).
Full build green: go build -tags 'cloud cloud_mount' ./...
FLAGGED for cutover review:
- ASSUMED table/column schema (worker source ships dist-only; reconcile
traces/observations/scores columns with 002_llm_observability.sql before the
console producer is repointed).
- Durable EmbeddedTasks hand-off: flush runs INLINE today (removes BullMQ+Valkey);
the durable enqueue->activity path is the next reviewed step (a Datastore insert
must be a durable Activity, not run in a workflow fn).
Pre-existing stray direct modernc.org/sqlite import in a test file (from #201).
Swap to the canonical _ "github.com/hanzoai/sqlite" (its !cgo backend IS modernc,
identical behavior) so NO file anywhere imports modernc directly. Test-only —
not in the shipped binary — but keeps 'hanzoai/sqlite only' truly airtight.
ONE data model, MANY adapters: handlers keep returning structured JSON; this
middleware re-serializes successful application/json responses through
zap-proto/md when the caller asks (Accept: text/markdown or ?format=md), so
token-efficient markdown is a request-time choice, not a second code path.
/v1/code/ + /v1/agents/ may default to markdown; caller override always wins.
Fail-safe: md render error leaves JSON untouched (never a 500); streams/HTML/
bytes pass through.
Native, per-org code-intelligence subsystem for AI coding agents and hanzo.app.
Retrieval is HYBRID — three orthogonal tiers fused with reciprocal-rank fusion
(the SOTA lesson that embeddings alone under-serve code search):
- lexical — FTS5 trigram over code-tokenized text (camelCase/snake_case split,
operators kept); substring + regex via trigram pre-filter + regexp verify (Zoekt).
- symbolic — go/parser for Go (real def→ref call edges) + compact lexical
extractors for TS/JS/Python/Rust/Solidity; go-to-symbol + edge table.
- semantic — AST-boundary chunks embedded via the SAME gateway /embeddings
clients/knowledge uses; cosine KNN over a float32 vector table.
Storage is ONE SQLite file per org at {DataDir}/orgs/{slug}/code.db (HIP-0302):
the tenant boundary is PHYSICAL. Every request is principal-gated (principal.Tenant)
— no validated principal ⇒ 403, a client X-Org-Id is never trusted.
Routes (order 134, before the AI /v1/* catch-all):
GET /v1/code/search ?q=&type=text|regex|symbol|semantic|hybrid&repo=&limit=
POST /v1/code/context {query,budgetTokens,repo} → budget-packed context bundle
GET /v1/code/ask ?q=&repo= (or POST) → cited RAG answer (deps.AI)
POST /v1/code/index {repo,files,prune} → (re)index, incremental by hash
Parsing is pure-Go and vectors are brute-force cosine because the repo's canonical
build is CGO_ENABLED=0 (Makefile); CGO tree-sitter and the sqlite-vec loadable
extension would break `go build ./...`. The vectors table is the schema-compatible
sqlite-vec `vec0` drop-in seam. Builds + tests green under both CGO=0 (modernc) and
CGO=1 (sqlite_purego).
listMachines/listGpus already log+fall-through to a BYO-only list when Visor is
down; listClusters alone returned the error, which surfaced as a 502 + a console
error on the Clusters/GPUs page for every org where Visor isn't deployed
(visor.hanzo.svc unresolvable). Mirror the graceful fold: log, drop managed
pools, still return the org's BYO clusters. 200, honest empty, no page error.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Unblocks the embedded IAM subsystem. v1.31.18 getPermissionEnforcer called
authz.NewEnforcer(&DefaultLogger{}, false) — the (logger,bool) form Casbin
type-switches params[1]->persist.Adapter, panicking "bool is not persist.Adapter:
missing method AddPolicy" the moment InitEmbed runs against a FRESH store (exactly
the unified cloud iam subsystem). v1.31.19 (fe50caf7) uses NewEnforcer()+SetLogger.
Proven on a fresh store: CLOUD_ENABLE=iam,ai boots green ("iam embedded in-process")
and serves /v1/iam/.well-known/openid-configuration 200 (was fail-closed 503/404);
iam + ai co-reside, process listens :8080/:9653/:9090.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Finishes the native web-search half of /v1/websearch (HIP-0106: no external
search SaaS, one fewer non-Go dependency). The route no longer reverse-proxies
to the retired SearXNG pod; it runs metaSearch in-process:
- search.go: keyless meta-search over public engines (Bing default, DDG opt-in),
parses HTML with x/net/html, merges+dedupes by normalized URL, returns the
exact SearXNG {query,number_of_results,results[]} envelope the LibreChat
searxng client decodes verbatim. A failing/bot-challenged engine contributes
zero and never fails the request (degrades to fewer results, never a 5xx).
- websearch.go: /v1/websearch/search now serves searchNative; removed
newSearchProxy + searchUpstream + WEBSEARCH_UPSTREAM. Auth unchanged (F2):
validated principal OR shared X-API-Key; neither ⇒ 401/503, never open.
- tests: converted the proxy route/guard tests to native (mocked engine via
WEBSEARCH_BING_URL fixture); added metaSearch parse + graceful-degrade tests.
go test ./clients/websearch/... green; full binary builds.
Wire the metrics board handler and route that the completed data layer was
missing. metrics.go already had the ClickHouse ledger + GenAI-span aggregation
(assembleTotals/Series/ByModel, usageWhere, latency percentiles) and the
in-memory honest-empty path; this adds:
- metricsBoard handler: principal gate (403 without a validated tenant),
SuperAdmin all-orgs via c.IsAdmin(), range preset -> window/bucket
(24h|7d|30d, ?interval override), ?project threaded. nil telemetry or a
non-default project -> honest-empty board (never 503, never fabricated).
- Telemetry.Metrics added to the interface (dsTelemetry + memTelemetry already
implement it).
- Route registered in Mount() and the test mountApp().
All clients/eval tests pass, including the three handler tests that previously
404d (TestMetricsHandlerHonestEmpty / RequiresPrincipal / NonDefaultProjectEmpty).
Co-authored-by: hanzo-dev <dev@hanzo.ai>
* feat(eval): drop cloud_usage-as-observations read; collapse to o11y span plane
Step 1 of the unified AI-observability plan: the observation of record is the
o11y gen_ai span plane (/v1/o11y/observations), not a second projection of the
metering warehouse. evalsvc no longer reads hanzo.cloud_usage as 'observations'.
- Remove Telemetry.ListObservations + Observation/ObservationFilter types + the
dsTelemetry (cloud_usage) and memTelemetry impls + the now-dead asInt64 coercer.
- Remove GET /v1/evals/observations route, its handler, and observationView/
toObservationView. The console Observe > Observations view now reads o11y.
- cloud_usage stays the metering warehouse (read by /v1/usage + billing) — only
the duplicate obs projection is gone. eval keeps its unique datasets/evaluators/
runs and its own eval_traces/eval_scores tables.
- Bump ai dep to the enriched-gen_ai-span build (forward-only from v1.804.0).
Build+vet+test ./clients/eval green.
* cloud(cli): compute ladder — hanzo run/agent/bot verbs (thin /v1 clients)
Preserve in-flight CLI work: hanzo run (artifact/function), hanzo agent
(headless managed agent -> /v1/agents/:ref/run), hanzo bot (computer-using
agent -> operative/visor). Thin clients over IAM token + cloud /v1.
Claude-Session: https://claude.ai/code/session_01EjRSpFBvbjxTqaYVbds9bA
world.hanzo.ai is the OIDC client `hanzo-world`; IAM stamps its access
tokens with aud=hanzo-world (each app's aud is its client_id). That
audience was missing from cloud's baked identity-sanitizer allowlist, so
signed-in world tokens resolved anonymous and api.hanzo.ai returned 401.
Append the client_id (forwards-only) and add a membership + resolved-env
acceptance test.
Claude-Session: https://claude.ai/code/session_013jh8aka8q8RvhhVQ1psMeW
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Cloud now serves /v1/commerce/* + /_/commerce/* ITSELF via a new
clients/commercesvc leaf that wraps commerce.Embed's gin handler — the same
wrap-don't-rewrite fold clients/iam and clients/kms use — instead of proxying to
a remote commerce pod. deps.Commerce resolves to the in-process CommerceClient
(CommerceInProcess) when commerce is co-resident.
Why a leaf (not the upstream commerce.Mount): commerce's own Mount/init sit
behind //go:build cloud, and cloud builds without that tag, so cloud's plain
build never compiled the registration — commerce was blank-imported yet absent
from cloud.Registry (proxied at runtime). The leaf lives in-repo and imports only
commerce's cloud-free surface (Embed, api.Route, Version), so the upstream
commerce.Mount that imports cloud stays tagged out — no cloud<->commerce import
cycle. Zero go.mod/go.sum churn (commerce v1.46.40 already required).
STAGED (prod-safe): commerce joins iam/ingress in stagedSubsystems, so the
mount-all default is unchanged — the in-process cutover happens only on explicit
CLOUD_ENABLE=...,commerce. Phase 2 flips the default once validated in prod. The
remote proxy seam (CLOUD_COMMERCE_HTTP_URL / CLOUD_COMMERCE_ZAP_ADDR) is
untouched; the disabled/RPC fallbacks still compile.
GetTenantConfig is answered in-process (org + brand); CheckEntitlement fails
closed until commerce exports its subscription->plan->features resolver (Phase 2)
— the specified "cannot verify => never open" default clients/entitlements relies
on.
Proven: CLOUD_ENABLE=commerce -> GET /v1/commerce/tenant, /v1/commerce/catalog,
/_/commerce/healthz all 200 from the embedded gin engine; an unknown
/v1/commerce path returns gin's own 404 (a proxy would 502).
Claude-Session: https://claude.ai/code/session_016yg7GPhYdWCh9vpp4HEwLZ
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Served sites typed .wasm/.data/.mem/.unityweb/.pck via mime.TypeByExtension only,
which returns "" for the engine payloads (empty Content-Type) and can mis-type
.wasm — breaking WebAssembly.instantiateStreaming (requires application/wasm) and
the loader's streaming fetch of Unity/Emscripten .data/.mem and Godot .pck. Pin
those in one gameAssetType map; everything else defers to the stdlib table.
Unblocks hosting WebGL game builds. Tested (TestGameAssetContentType).
'arcd' is the client-side GitHub-Actions BYO product (github.com/arc-runner);
the platform's own native CI/compute pool is 'runner'. Rename the enqueue path
and de-brand the build command help/comments accordingly. Pairs with the
platform route move pages/api/v1/arcd/enqueue.ts -> pages/api/v1/runner.ts.
The 4 go invocations (mod download, sqlite double-register gate, encryption-proof
test, final build) had no cache mount, so every release recompiled the full CGO
graph (k8s + otel-collector + sqlcipher, CGO=1) from scratch on the ephemeral ARC
runner — the Build step was ~1316s/22min, dwarfing every other step. Mount
/go/pkg/mod + /root/.cache/go-build (sharing=locked) so the persistent ARC dind
BuildKit cache keeps the Go build+module cache warm. First build cold; subsequent
builds reuse compiled artifacts — target single-digit-minute rebuilds.
CreateTraces opens the ClickHouse conn + spawns the writer's ticker goroutine, so
a failed Start must Shutdown to release them rather than leak on the fail-soft
mount path.
The validated-org-principal check asserted ==200, coupling this cross-package
gate test to the orgs/users/me handler's downstream success (IAM/datastore),
which flaked under full-suite parallel load. Assert the gate/shadow decision
only — admitted (not 403) and mounted (not 404) — since tenant-scoped data
correctness is proven in clients/admin/scope_test.go. Anonymous->403 and the
SuperAdmin-only platform routes are unchanged.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The prior two attempts (c97af12 sha-pin via git ls-remote, 4a7e533 via gh api)
both FAILED at version-compute: the ARC runner has no gh CLI, and git ls-remote
404s because actions/checkout installs a global http.extraheader carrying THIS
repo's GITHUB_TOKEN (scoped to hanzoai/cloud), overriding URL creds on the
cross-repo hanzoai/console lookup.
Bulletproof instead: no console-HEAD resolution at all. release.yml passes the
cloud commit sha as --build-arg CONSOLE_CACHEBUST (unique per push); the
Dockerfile references it in the proven `git clone --depth 1 --branch main`
RUN, so the layer cache key changes every build and re-clones console main HEAD
fresh. No gh, no ls-remote, no extraheader. Correctness over cache reuse — the
console stage rebuilds each release, but the embed is never the frozen snapshot.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- clients/base: register /v1/base/health in Mount BEFORE the CLOUD_BASE_EMBED
gate and mark cloud.HealthOwner — same always-on liveness pattern as
clients/plan + clients/pricing. Fixes cmd/cloud TestMountAllAndServeHealth
(base was the only listed subsystem not self-serving health).
- clients/kmssvc red_dualmount_test: #192 made the admin cockpit two-scope —
orgs/users/me are org-scoped (guardScoped: a validated org principal is
admitted + hard-scoped to its own org; anonymous still 403), while
audit/roles/finance/flags/revenue stay SuperAdmin-only (guard: 403 for a
non-admin principal). Assert both, plus kms's public /v1/kms/config never
shadows either. No production code semantics changed — the stale test tracked
the pre-two-scope admin-only contract.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
git ls-remote failed 'Repository not found' on the cross-repo hanzoai/console
lookup: actions/checkout installs a global git http.extraheader carrying THIS
repo's GITHUB_TOKEN (scoped to hanzoai/cloud only), which overrides the URL
creds and 404s. gh api honors GH_PAT (org read) and is unaffected. Unblocks the
console-embed-freshness fix (c97af12) — the release aborted at version-compute
before building.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
cloud already embeds the o11y trace write side (chtraces), so shipping its OWN
spans over the ZAP wire to a collector that then writes the same store is pure
waste. Route them through the ZAP locality-adaptive Router (luxfi/zap v1.2.1):
when sender and sink share this binary, the Cost-0 InProcessInterface wins and
the LIVE proto batch is handed to the sink by value — zero ZAP-wire serialize,
zero socket, no second collector hop.
- clients/o11y/tracesink.go: Router + Cost-0 InProcessInterface on Destination
"hanzo.o11y.traces"; a chtraces exporter (the REAL o11y_index_v3 writer, reused
as a consumer.Traces — its pdata->SpanV3 conversion is unexported) fed in
process. Handler bridges SDK-exporter proto spans -> pdata via one in-memory
OTLP round-trip. OPT-IN (O11Y_TRACES_ZAP_INPROCESS) + fail-soft: any error
leaves cloud's spans on the wire; can never take cloud down. NewTraceExporter +
routerTraceClient own the transport; cmd/cloud stays the composition root.
- cmd/cloud/telemetry.go: install the ONE tracer provider over the Router
(in-process primary, ZAP wire fallback when the sink isn't registered). Enable
when the in-process sink is on OR a wire endpoint is set. Composition-root
single-provider invariant (ai's GenAI tracer inherits it) preserved.
- go.mod: github.com/luxfi/zap v1.2.0 -> v1.2.1 (adds Router/InProcessInterface).
TDD: span from cloud's provider reaches the in-process handler with no wire
client and no socket; proto->pdata round-trip preserves the span; router prefers
in-process, falls back to wire on ErrNoRoute, surfaces ErrNoRoute when neither.
The kmssvc dir was an artificial split: clients/kms is the KMS library
(embeds luxfi/kms + SecretStore + in-process client), clients/kmssvc was
the Fiber subsystem mounting /v1/kms/* — and it was ALSO package kms, dir-
named kmssvc only to dodge a dir-name collision. That svc suffix is a
workaround, not a concept.
Move every kmssvc file into clients/kms (kmssvc.go → mount.go; login.go,
env_required/kms/login_ratelimit/paas_sync/red_*/v6 tests). subsystems.go
imports clients/kms (order 10, /v1/kms/*); exactly one cloud.Register("kms").
Zero kmssvc refs remain. Full cloud build + kms package tests (incl red_*
adversarial) green. (Pre-existing clients/kms/replication test build failure
is unrelated — broken on main before this change.)
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Composes the ha lease-round (v0.1.1) and vfs/replica.FencedStore (v0.6.3) into
the cloud per-org substrate, and adds the request-layer exactly-once dedup, so
per-org SQLite is safe under a multi-replica Deployment (not just replicas:1).
Four orthogonal concerns, one home each:
internal/org/fence.go CASFencer: the INTERIM monotone round source. A per-org
writer lease {round,owner} over the object store's CAS
(a single linearizable register); takeover strictly
bumps the round, renewal keeps it. Implements ha.Fencer,
so the Lux BFT round drops in behind the same seam later.
HRW is only an optimization (cuts contention); safety
does not depend on a fresh/agreed membership view — a
split view costs liveness, never safety, because the
fence backstops it.
internal/org/condstore.go MinioConditionalStore: the concrete atomic-CAS store
(minio If-Match against the SeaweedFS gateway), promoted
from the orphaned internal/writefence. satisfies
replica.ConditionalStore.
internal/idem/ exactly-once request execution: request-id PK written in
the SAME txn as the effect (atomic dedup+effect), shipped
in the per-org snapshot so a retry re-routed after a
rolling upgrade is deduped on the successor. 'fail if
already done' via ErrAlreadyApplied.
internal/org/shared.go re-export FencedStore/ConditionalStore/Lease/Fencer/
Round/ErrStaleRound so the org API stays one surface.
Deletes internal/writefence (was orphaned, zero importers): its fence primitive
is promoted to vfs/replica.FencedStore (the storage substrate's rightful home),
its minio store to condstore.go — one and one way, forward-only.
Safety (no data loss + no double-exec) rests on the composition, proven by
handoff_test.go against the four hazards: (a) partition minority cannot advance
the round -> cannot write; (b) rolling-upgrade handoff -> successor CarryForwards
the predecessor's last landed write + dedups; (c) duplicate request -> idem runs
once; (d) deposed writer -> refused by election, and if it still ships, fenced by
FencedStore. Ship-before-ack: a request is 'done' only once its fenced ship lands.
Interim round source = single linearizable register (object CAS) = crash-fault
tolerant. Roadmap: replace readLease/claim with the quasar PQ-BFT agreed round
(Byzantine-tolerant, deterministic finality, 3/5 quorum) — same ha.Fencer seam,
same FencedStore admission.
Cross-repo: pins ha@fence-lease + vfs@fenced-store (pseudo-versions); retag to
ha v0.1.1 + vfs v0.6.3 once those merge. go.sum touches only ha+vfs; the
pre-existing luxfi/keys@v1.2.0 re-tag mismatch (blocks `go mod tidy` on main
today) is unrelated and untouched.
The console-clone+build layer was keyed only on static text ('git clone
--branch main'), so on the persistent ARC dind BuildKit cache EVERY cloud
build re-embedded the SAME stale console snapshot. New console work — the
native Tracker module, and everything since the cache was first warmed —
silently never shipped: console.hanzo.ai/tracker rendered an old surface
with zero /v1/tracker calls even on a freshly-deployed image.
Fix (values, not places): release.yml resolves hanzoai/console main HEAD
(git ls-remote) at build time and threads it through --build-arg CONSOLE_REF;
the Dockerfile fetches that exact ref (init+fetch+checkout, sha- or branch-
capable). A changed sha moves the layer cache key, so each build embeds the
live console commit — deterministically pinned, never frozen.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
admin.hanzo.ai as ONE cockpit for BOTH tiers off ONE identity predicate. scope.go decomplects the rule into a single place (resolveScope/scopedOrgs/descendants): owner==admin (c.IsAdmin, SuperAdmin) => cross-tenant, all orgs; any other validated admin => their OWN org, hard-scoped server-side. guardScoped admits a SuperAdmin OR a validated org-pinned caller and the handler scopes the data; the platform control plane (roles/audit/finance/revenue + launch/release/flags/access) stays s.guard (SuperAdmin only). me/overview/orgs/users/usage/analytics/bases are org-scoped; a non-super caller can never read another tenant for any input (their org is the sanitized, un-forgeable c.Org()).
Feature flags / launch / access via Hanzo Insights (one engine, not two): clients/featureflags is a hot-apply evaluation seam over Insights /flags (env = fallback default, 15s TTL, fail-safe degrade); /v1/admin/flags surfaces the launch switches (public_signup, waitlist_open, waitlist_access_capacity, ...) with deep-links to the Insights flag manager + activity log. /v1/admin/waitlist + /boost proxy the Base waitlist engine (server-authed, KMS secret, audited grant). /v1/admin/bases is the scoped tenant-Base panel seam (honest-empty until the Base engine is embedded).
Fix: waitlist.go shadowed the ok() envelope writer with a local bool (compile error) — renamed to configured. Tests: scope_test.go proves the two-scope invariant (super sees all; org-admin hard-pinned to own org; platform routes 403 an org-admin; users read pinned to own org); featureflags_test.go proves hot-apply + env fallback. go build ./... green; go test ./clients/admin/ + ./clients/featureflags/ green.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
* feat(edge): embed the gateway CORS + per-IP rate-limit role in cloud
Fold the hanzoai/gateway edge role into cloud so it can serve api.hanzo.ai
directly, dropping the redundant KrakenD hop. cloud already validates the IAM
JWT + strips/re-mints identity headers (SanitizeIdentity), runs the per-tenant
ScopeRateLimit, and owns balance/spend-cap quota (BillingGate) — the gateway
duplicated the JWT/identity role and added only CORS + a per-IP flood cap.
Adds middleware_edge.go (package cloud), two orthogonal middlewares wired into
the serve.go chain AFTER Logger/sites and BEFORE identity:
- EdgeCORS: credentialed reflect-Origin CORS matching the gateway's policy
(methods/headers/max-age). DEFAULT OFF (empty CLOUD_CORS_ORIGINS) so the
shared Traefik ingress `cors-allow-all` stays the sole CORS authority on the
recommended rollout — enabling both would double the ACAO header. Set
CLOUD_CORS_ORIGINS only on a direct DO-LB->cloud edge. Handles + short-circuits
the OPTIONS preflight (204) before any auth work.
- EdgeRateLimit: per-client-IP fixed-window flood cap (default 100/1s, gateway
service-tier parity, strategy:ip) that runs BEFORE identity — the one gap
ScopeRateLimit (keyed on the validated tenant) structurally can't see: an
anonymous flood with no valid JWT. Keyed on the leftmost X-Forwarded-For;
in-cluster direct callers (no XFF) are exempt, matching the standalone
gateway's public-only scope. Opportunistic eviction keeps the bucket map
bounded at edge IP cardinality (default ON, CLOUD_EDGE_RATELIMIT=false to
disable). This preserves the gateway's protection rather than dropping it.
Config: CORSOrigins, EdgeRateEnabled/PerIP/WindowSec (config.go).
Tests: TestOriginMatcher, TestEdgeCORS_*, TestEdgeRateLimit_* (all green).
CGO_ENABLED=0 go build ./... + go test . green.
* feat(gateway): /v1/gateway runtime-mutable edge-policy plane
Make the embedded gateway role RUNTIME-CONFIGURABLE instead of baked into static
config: an operator/SuperAdmin retunes CORS, the per-IP flood cap, or a tenant's
rate ceiling via PUT /v1/gateway/config with NO redeploy — replacing the gateway's
image-baked KrakenD config.
clients/gatewaypolicy (leaf pkg, stdlib + hanzoai/sqlite only, no cloud import so
both the middleware and the HTTP subsystem share it cycle-free):
- Policy{CORSOrigins, PerIPRPM, WindowSec (platform), OrgRPM (per-org)} — every
field is enforced by a consumer; no stored-but-ignored knob.
- Store: one encrypted per-tenant SQLite (gateway.db), org-keyed rows. The admin
org row is the PLATFORM policy, layered over the static env/flag boot defaults.
Cached resolvers Platform()/OrgRPM()/Effective() (5s TTL, fail-open to static),
merge(base,over) makes a partial PUT additive. Fail-soft: a store-open error
degrades to static-only (reads work, writes error) — the edge never goes down.
clients/gatewaysvc: the /v1/gateway subsystem (order 139) — GET/PUT config over
the SAME store, IAM-gated like clients/pricing/enablement.go:
- platform fields (CORS/per-IP) writable ONLY by SuperAdmin (c.IsAdmin()); routed
to the platform row explicitly (PutPlatform) so an org-switched SuperAdmin still
lands on it.
- per-org OrgRPM writable by the org admin (own org via principal.Tenant, never a
raw header) or a SuperAdmin targeting ?org=<slug>.
Wiring: deps.GatewayPolicy (BuildDeps constructs it, layered over staticEdgePolicy;
serve.go closes it at shutdown). EdgeCORS/EdgeRateLimit now read the PLATFORM policy
LIVE (recompiling the CORS matcher only when the allowlist changes; the per-IP
limit/window per request). ScopeRateLimit gains the runtime per-org OrgRPM
override (most-restrictive-wins with the commerce-configured ceiling).
Tests: gatewaypolicy (static-only, platform layering, additive merge, per-org +
platform-default OrgRPM, persist-across-reopen); gatewaysvc (principal required,
org-self OrgRPM, org-admin platform 403, SuperAdmin platform, org-switched-still-
platform, empty-body 400). CGO_ENABLED=0 go build ./... + go test . green.
Mounts /v1/waitlist/* served in-process off the embedded hanzoai/base app over
the durable cloud PVC — the in-binary replacement for the standalone superbase
pod. STAGED + fail-closed: no-op unless CLOUD_BASE_EMBED=1. Registered as the
"base" subsystem (order 60) in clients/base.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The datastore connects ASYNCHRONOUSLY: ai/object.InitDatastore flips
DatastoreEnabled true only AFTER Mount returns. Mount ran the CREATE TABLE
DDL only when DatastoreEnabled() was already true, so in prod it was skipped
and never retried -> GET /v1/sbom/{ref} 502 'Unknown table expression
identifier hanzo.sbom_component' while /v1/sbom/health reported datastore:true.
Add a lazy, idempotent ensureTable(ctx) guarded by a mutex+bool that latches
ONLY success (a transient failure retries; sync.Once would cache the failure).
It CREATE DATABASE IF NOT EXISTS hanzo then CREATE TABLE IF NOT EXISTS, and is
called from ingest and resolve right after requireDatastore() passes; on error
they return a retryable 503. Mount now routes its best-effort boot DDL through
the same ensureTable and is non-fatal (a Mount-time miss no longer aborts the
subsystem).
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The live cloud console runs on <brand>.cloud hosts that route straight to the
cloud Service (console.lux.cloud, console.zoo.cloud, …). BrandForHostOK only
matched a brand's primary marketing Domain (lux.network, zoo.ngo), so a request
Host on lux.cloud/zoo.cloud fell through to the deployment brand — emitting Hanzo
branding on a Lux/Zoo surface (agent-skills catalogue + any Host-branded reply).
Add AltDomains per brand (lux→lux.cloud; zoo→zoo.network,zoo.cloud;
hanzo→hanzo.cloud,hanzo.app; pars→pars.ai) and match them in BrandForHostOK.
Base-URL/issuer scoping still uses the primary Domain.
Claude-Session: https://claude.ai/code/session_01CDooqWJiB7yNNaSjGQjdL7
Co-authored-by: hanzo-dev <dev@hanzo.ai>
New subsystem clients/agentskills serves the Agent Skills Discovery surface from
a catalog GENERATED by hanzoai/openapi's skills.py and embedded via go:embed:
GET /.well-known/agent-skills/index.json the brand's MASTER catalogue
GET /.well-known/agent-skills/<skill>/SKILL.md one skill document
WHITE-LABEL: the brand is decided per request from the Host (new BrandForHostOK,
mirroring platform.ts getWhiteLabelBrand) — api.hanzo.ai serves Hanzo,
api.lux.network serves Lux (lux.id), api.zoo.ngo serves Zoo — never Hanzo
branding on a Lux/Zoo surface. An unmatched Host degrades to the deployment brand
(CLOUD_BRAND), not blindly Hanzo. Order 8 registers these exact routes BEFORE
IAM's /.well-known/* wildcard (50) and the console catch-all, so they win Fiber's
first-match. Public, GET-only, no secrets.
The binary does not re-derive skills — it serves the embedded bytes, so the
sha256 digests in index.json match the served SKILL.md exactly. Only a tiny,
self-consistent `ai` fallback is committed (catalog/.gitignore); `make
agentskills` / the Dockerfile `skills` stage regenerate the FULL catalog (all 68
services × hanzo/lux/zoo) from the openapi SOT before `go build`, mirroring
webui/dist. End-to-end serve test drives the real router (index + SKILL.md +
white-label + digest + 404); cmd/cloud links clean.
Claude-Session: https://claude.ai/code/session_01CDooqWJiB7yNNaSjGQjdL7
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Add clients/sbom: a self-contained subsystem riding the ONE shared
ClickHouse client (ai/object.Datastore*) — no second connection — that
ingests CycloneDX SBOMs from CI and serves them by image digest/ref.
The store (hanzo.sbom_component, ReplacingMergeTree) is GLOBAL/cross-tenant
by design: an SBOM belongs to a content-addressed image digest, not a
tenant, so any tenant deploying that image resolves the same component set.
Ingest is gated to the canonical cloud super-admin check (c.IsAdmin(),
owner==AdminOrg) which the build fleet carries; resolve exposes only an
image's immutable bill-of-materials (no tenant data).
POST /v1/sbom ingest (super-admin/CI): flatten document.components[]
GET /v1/sbom/{ref} resolve by digest OR ref (FINAL dedupe, type,name order)
GET /v1/sbom/health liveness + datastore bool (not JWT-gated)
Registered id "sbom" order 137 with cloud.HealthOwner (binds before the ai
/v1/* catch-all at 150). Mirrors clients/analytics for structure, coercers,
and the honest-503 datastore gate.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Clean-semver pin bringing three ai releases into the cloud binary:
- v1.803.1 fix(account): cookie-session self-heal (#71) — a signed-in admin
whose beego session already holds a guest u-<hash> is rebound to the
canonical identity from hanzo_iam_token, so /v1/admin/* stops 403ing.
- v1.804.0 refactor(authz): ONE super-admin rule — membership in the `admin`
org (owner == AdminOrg); drops the configurable globalAdminOrgs + built-in.
Matches cloud's clients/admin (isSuperAdmin canonical, isGlobalAdmin alias)
and the console isSuperAdminAccount gate.
Supersedes the pseudo-version pin (#197) and the intermediate v1.803.1 pin
(#198, closed). Verified: go build -tags "libsqlite3 sqlite_fts5" ./cmd/cloud.
Claude-Session: https://claude.ai/code/session_01VZbTTNtf4y8y3XUr9wMqTX
Co-authored-by: hanzo-dev <dev@hanzo.ai>
* feat(platform,sites): pure-Go zip/tar.gz static-site upload + custom-domain serving
Adds a self-service static-site deploy to the unified cloud binary's PaaS
surface and lets the site edge serve a customer's own domain from S3.
- projects: walkArtifact accepts a ZIP (archive/zip) as well as tar(.gz),
sniffed by magic bytes; one deploy contract (index.html at root, same size
and traversal guards), three container formats. A single wrapping top-level
directory (a zip made from a project folder) is stripped so index.html lands
at the root.
- projects: the deploy handler reads the artifact from a multipart file upload
(a browser <input type=file>) OR the raw request body (a curl one-liner).
- sites: the host edge now serves a bound CUSTOM domain (a customer apex/host
pointed at this edge) from that project's S3 prefix, resolved by the full
host. Only external hosts (never one of our self domains) with a LIVE binding
are served; every other host — our api/console hosts, or an unbound host
routed here — Continues to the normal pipeline, so the API path pays no
per-request lookup and a customer binding can never shadow a real Hanzo host.
- projects: POST/GET .../domains binds and lists a site's custom domains
(admin-gated until DNS-ownership verification is wired here).
- surface: the static engine is exposed under /v1/platform/sites/* (the PaaS
namespace) in addition to /v1/projects/*, so the one user flow is create a
site -> upload a zip -> bind a domain -> live.
Pure Go, CGO-off. New unit tests cover the zip walker, format dispatch,
single-root strip, custom-domain routing (served/passthrough/self-host/not
-live), hostname validation, and host binding.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(projects): authorize custom-domain binding by the platform-operator org
A custom-domain bind is authorized for a global admin OR the platform-operator
org (the deployment's brand org, env CLOUD_PLATFORM_OPERATOR_ORGS, default the
brand). The operator manages customer DNS until per-tenant DNS-ownership
verification is wired here. Safe because a bound domain is inert until its owner
points DNS at this edge — the real gate is DNS control.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* debrand: signoz -> o11y (branding) + repoint collector module
Drop "SigNoz"/"signoz" where it is BRANDING (comments, docs, prose) to o11y,
and repoint cloud's direct collector import to the renamed module.
- go.mod/go.sum: github.com/hanzoai/signoz-otel-collector v0.144.6 (direct)
-> github.com/hanzoai/otel-collector v0.144.7 (direct). The old module stays
as an // indirect dep because hanzoai/o11y v1.5.2 (separate repo, out of
scope) still imports it; the fork also pulls upstream
github.com/SigNoz/signoz-otel-collector v0.144.5 // indirect.
- clients/o11y/ingest.go: chlogs/chtraces imports -> otel-collector.
- Branding prose swapped to o11y in telemetry.go, subsystems.go, embed.go,
logs.go, metricsread.go, scope.go, ingest_test.go, agents.go, LLM.md,
docs/consolidation.md.
KEPT (not branding):
- ClickHouse schema read by cloud (written by the deployed collector):
signoz_traces / signoz_logs / distributed_signoz_index_v3, severity_text
columns. Renaming reads without migrating the live schema breaks them; the
v0.144.6->.7 patch bump does not migrate table names.
- Upstream API in hanzoai/o11y/pkg/signoz: import path, alias, type SigNoz,
and signoz.New / SigNoz.Start references (that repo is debranded separately).
- Upstream package names: signozclickhousemetrics.
- Honest attribution: "SigNoz's dd-sketch fork of ch-go".
Build: go build ./... = 0, go vet = 0, go test ./clients/o11y ./clients/admin
= ok. go mod tidy is blocked by a pre-existing luxfi/keys@v1.2.0 go.sum
checksum mismatch (identical on origin/main) -> CI-authoritative.
* o11y: read o11y_* ClickHouse tables + bump collector to v0.144.8
Direct ClickHouse table reads renamed signoz_* -> o11y_* to match the
o11y read plane and the data-preserving RENAME migration (hanzoai/o11y#28):
o11y_traces.distributed_o11y_index_v3 (was signoz_traces.distributed_signoz_index_v3)
o11y_logs.distributed_logs_v2 (was signoz_logs.distributed_logs_v2)
Files: clients/o11y/logs.go, clients/o11y/metricsread.go,
clients/o11y/ingest.go, clients/admin/o11y.go (+ o11y_test.go).
Bump github.com/hanzoai/otel-collector v0.144.7 -> v0.144.8 (writer side
now CREATEs/WRITEs the same o11y_* physical schema). Collector go.mod is
unchanged between the two tags (identical go.mod hash) — pure source
rename, so the module graph is unchanged; go mod tidy left to CI
(pre-existing luxfi/keys tidy block is CI-authoritative).
go build ./... = 0. clients/admin + clients/o11y tests green (SQL
assertions now match o11y_* target names). Lockstep deploy: collector
v0.144.8 -> o11y#28 RENAME migration -> o11y+cloud readers.
* cloud: embed o11y v1.5.4 — version-less /v1/o11y + o11y_ schema reads + debrand
Bumps hanzoai/o11y v1.5.2→v1.5.4 (version-less surface + o11y_ ClickHouse table
reads + the signoz→o11y debrand) and repoints embed.go to the renamed runtime
package pkg/signoz→pkg/o11y (type SigNoz→O11y). Pairs with otel-collector v0.144.8
(writes o11y_) + the lockstep cutover migration. go build ./... = 0.
---------
Co-authored-by: hanzo <z@hanzo.ai>
Embeds hanzoai/o11y#26: the mount normalizes the /v1/o11y/<resource> public
contract onto the internal SigNoz /api/vN routes (kills the /api/ leak, fixes
the llmobs /v1/o11y/* 404). Pairs with the cloud CR O11Y_GLOBAL_EXTERNAL__URL=""
change (universe#461) — deploy together.
Co-authored-by: hanzo <z@hanzo.ai>
The ai (casibase) layer SETS the httpOnly hanzo_iam_token cookie (the IAM JWT) after
login (ai/controllers/account.go iamTokenCookieName), but cloud's own identity
middleware only read [iam_access_token, access_token, hanzo_token] — NOT
hanzo_iam_token. So the embedded console (browser holds ONLY that cookie, no
Authorization header) resolved to no validated principal → every org-scoped /v1
endpoint (agents, gpus, machines, platform, orgs, entitlements, …) 403'd
'X-Org-Id required', and modules rendered empty. Add hanzo_iam_token (first) to
cookieTokenNames so cloud reads the SAME cookie the ai layer sets → validates the
JWT → X-Org-Id from owner → org-scoped surfaces authorize. Verified: the JWT is
present in the browser (1533-char httpOnly hanzo_iam_token); only the name was wrong.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pulls the get-account cookie-path fix (hanzoai/ai#81): a signed-in admin
whose beego session already holds an anonymous guest u-<hash> is now
self-healed from the hanzo_iam_token credential to its canonical identity,
so /v1/admin/* stops 403ing under the console cookie session. Verified:
cloud binary builds with -tags "libsqlite3 sqlite_fts5".
Claude-Session: https://claude.ai/code/session_01VZbTTNtf4y8y3XUr9wMqTX
Co-authored-by: hanzo-dev <dev@hanzo.ai>
internal/org/shared.go split its aliases along the real seam: election
(Member/Owner/IsOwner/Replicas) now re-exports github.com/hanzoai/ha; the
Replicator/Store/DB/DBPath stay github.com/hanzoai/vfs/replica. membership.go
and cipher.go are unchanged (the alias types line up: org.Member = ha.Member).
writefence doc updated to name ha as the election primitive.
One concern, one home: who-writes (ha) vs how-state-ships (vfs). No behavior
change; internal/org + writefence pass with -race.
NOTE: `go mod tidy` is blocked in this repo by a PRE-EXISTING, unrelated
luxfi/keys@v1.2.0 go.sum checksum mismatch; ha was added via `go get` +
marked direct by hand. Re-run tidy once that pin is fixed.
Co-authored-by: hanzo <z@hanzo.ai>
Per RED's conditional GO on the encryption image:
- GOFLAGS -mod=mod -> -mod=readonly: the committed go.sum is the SOLE source
of truth; any needed-hash drift FAILS the build instead of silently
re-recording an unverified hash. Verified go.sum is complete + sumdb-
consistent (go build/download clean with GOSUMDB ON).
- Drop GOSUMDB=off: a money image must not blanket-disable the checksum
database. GONOSUMDB scoped to zap-proto/* only (first-party-direct).
- Digest-pin the three base images (node:24-alpine, golang:1.26-alpine3.22,
alpine:3.22) @sha256 for a reproducible money image.
The ldd/readelf link proof stays belt-and-suspenders behind the ciphertext
proof (verified at build time on the musl image).
The limiter guards the OFF-GATEWAY path where nothing trusted stamps
X-Forwarded-For, so keying on the client-settable XFF let an attacker send a fresh
value per request and reset the 30/min bucket at will. These are all post-auth
money-write routes, so key on the un-spoofable VALIDATED principal (X-Org-Id/
X-User-Id, minted by SanitizeIdentity from a verified JWT); fall back to the socket
peer (L4 RemoteAddr) for an unauthenticated request (which the handler 403s anyway).
Test now rotates XFF during the flood (must NOT reset) and asserts a second principal
keeps its own bucket.
Cloud-direct/off-gateway money path loses the edge WAF/limiter and the embed
session-bridge's Sec-Fetch-Site gate passes VACUOUSLY when Origin/Referer/SFS are
all absent (RED). Adds two positive controls to the console write surface:
CSRF (csrf.go) — GET /v1/console/csrf issues a token bound to the validated
principal (X-User-Id+X-Org-Id), MAC'd with keyed BLAKE3 (luxfi/crypto,
blake3.KeyedHash) under a server-only KMS key (CONSOLE_CSRF_KEY; ephemeral
per-process fallback). requireCSRF enforces X-CSRF-Token on the AMBIENT-cookie path
ONLY (no Authorization + a Cookie present) — Bearer/Basic/gateway/API callers are
immune to CSRF and skip it, so nothing non-browser breaks. A cross-site page cannot
read the same-origin token (SOP) nor set the custom header (no CORS preflight
granted), and the token is identity-bound so it can't be replayed as another user.
Rate limit (ratelimit.go) — per-IP token bucket (30/min) on mint/rotate/revoke key
+ wallet top-up; distinct from commerce's spend-cap, restores frequency protection
lost off-gateway. Keyed on XFF first-hop.
Wraps POST/DELETE keys, POST onboard, POST topup, POST billing, POST/PUT/PATCH/DELETE
commerce. Reads stay open. Tests: ambient-no-token 403, valid-token 200,
cross-identity replay 403, Bearer-skips-CSRF 200, rate-limit 429; existing suites
unchanged (no cookie ⇒ CSRF skipped). luxfi/crypto for the MAC (NOT stdlib/JWT).
Needs arcd build + RED review; CONSOLE_CSRF_KEY to be provisioned via KMS for
restart/multi-replica-stable tokens (coordinate ac742480).
The in-binary direct-Bearer path (console SPA -> cloud, gateway bypassed) stamps
X-User-Id = the JWT subject (a UUID) via idClaims.userID(). The console key ops
built the IAM id as <owner>/<X-User-Id> = hanzo/<uuid>, but IAM's mint-user-keys /
get-user resolve only <owner>/<name> (hanzo/z) -> 'password or code is incorrect'
-> hk- mint 502 on the cloud-direct path. (The gateway path worked because the
gateway minted X-User-Id == username.)
Fix (narrow blast radius, per RED-preferred approach — does NOT reorder userID()):
- idClaims.username(): the IAM username (name claim, then preferred_username),
NEVER the subject.
- SanitizeIdentity stamps X-User-Name from the validated username, DISTINCT from
X-User-Id. X-User-Name is already an authorityHeader (stripped on ingress,
re-injected only from validated claims -> forgery-proof).
- resolveCaller carries a distinct caller.username (X-User-Name, fallback to
X-User-Id for the gateway path); new caller.keyID() = <owner>/<username> is used
ONLY by the user-key ops. caller.id / caller.name are UNCHANGED, so the
billing/topup/commerce subjects are byte-identical -> zero money-path impact.
Tests: TestKeys_DirectBearerPath_MintsByUsernameNotUUID (mint targets hanzo/z, not
the UUID), TestSanitizeIdentity_StampsUserName, TestSanitizeIdentity_UserNameForgeryStripped;
existing key + identity suites unchanged (gateway path falls back to owner/name).
Needs arcd/CI image build (no local docker) + RED review before deploy.
The unified binary embeds IAM's per-org SQLCipher store and commerce's
per-tenant money DBs; the prior CGO_ENABLED=0 build shipped pure-Go
modernc — PLAINTEXT at rest. Rebuild CGO=1 against system libsqlcipher
(hanzoai/iam's proven recipe: libsqlite3 tag + libsqlcipher symlink +
-DSQLITE_HAS_CODEC), runtime base scratch -> alpine:3.22 + sqlcipher-libs
(CGO needs libc + the codec .so).
Baked-in RED gates (a failing gate = NO image):
- modernc double-registration guard: 0 modernc in the CGO=1 ./cmd/cloud
graph (the one 'sqlite' driver is mattn/SQLCipher).
- TestEncryptionProof: real ciphertext-at-rest under SQLITE_REQUIRE_CODEC=1.
- cek.go golden-vector KAT (TestUnwrapGoldenFixture + round-trip): a frozen
pre-luxfi-swap 61-byte DEK sidecar still decrypts under the shipped
luxfi/crypto-AEAD code — existing encrypted stores stay readable.
- readelf/ldd link proof: the binary binds sqlite3_* to libsqlcipher, never
a plaintext libsqlite3.
Console embed stage unchanged (same-origin console). RED must review before
the image ships.
Bump the six hanzo modules to their driver-converged releases so the CGO=1
unified binary has EXACTLY ONE database/sql 'sqlite' registration
(mattn/SQLCipher), ending the 'sql: Register called twice for driver
sqlite' panic:
sqlite v0.1.5 -> v0.2.3 (SetPersistWAL + OpenPragma primitives)
orm v0.5.2 -> v0.6.1
base v1.4.6 -> v1.5.7 (+ replicate v0.9.5, the last modernc leak)
commerce v1.42.29 -> v1.46.40 (+ go:embed plans fix)
o11y (pseudo) -> v1.5.1
replicate v0.8.0 -> v0.9.5
Retarget the mattn v2.0.3+incompatible replace v1.14.16 -> v1.14.47 (the
SetFileControlInt/SQLITE_FCNTL_PERSIST_WAL-capable version hanzoai/sqlite
v0.2.3 needs for SetPersistWAL).
Verified CGO=1: 0 modernc in the ./cmd/cloud dep graph; the 517MB binary
builds and boots (--help) with NO double-register panic.
studio.render now (1) writes any inputs shipped with the job into the
local studio input dir via its own /upload/image, so an uploaded photo
(which lives in orgs/{org}/input on the dispatching pod, unreadable here)
resolves for LoadImage before the render; and (2) forwards the job's
active org as the studio_active_org cookie on /upload/output, so the
finished render lands in that org's gallery even when the worker token's
home org differs (a@hanzo.ai home=hanzo, rendering for karma).
SuperAdmin: /v1/admin/me and /v1/admin/users now emit the canonical
`isSuperAdmin` key alongside the deprecated back-compat alias `isGlobalAdmin`
(both populated with the SAME fact — owner == AdminOrg). The console may read
either during the rename migration and sees the same truth. No DB change: the
signal was always a derived boolean, never a stored column.
Entitlements: new clients/entitlements subsystem (order 139) — the per-org
product-enablement plane the console's paid-product sidebar reads.
GET /v1/orgs/:org/entitlements -> { "enabled": [...] }
POST /v1/orgs/:org/entitlements { add?, remove? } -> { "enabled": [...] }
Two authorities, never braided: ENABLEMENT (this store: durable per-tenant
SQLite, (org,product) key, settings-store discipline) vs ENTITLEMENT (commerce:
deps.Commerce.CheckEntitlement at write time). A non-super-admin may only enable
a product the org's plan already grants (402 otherwise); disabling is never
gated; a super admin bypasses the commerce gate and may target any :org. Org
scoping mirrors clients/kms: :org must equal the validated owner claim unless the
caller is a super admin; a bearer-less forge fails the principal gate (403).
Tests (TDD, all green): store tenant-isolation + all-or-nothing Apply;
forged-request 403; cross-org 403; malformed org/product 400 (commerce not
consulted); entitled enable 200; unentitled enable 402 (nothing persisted);
super-admin bypass 200 (commerce not consulted); nil-commerce member-add 503
(fail-closed); remove never gated; empty mutation 400. Plus admin_test asserts
isSuperAdmin present and equal to isGlobalAdmin on both /me and /users.
Bumps @hanzo/plans to the World-pricing catalog (world-enterprise tier +
world.model_api gate) and adds the single-sourced enforcement contract for
the /v1/world data plane.
- clients/plan: export Entitlements(ctx, id) — the one Go seam to read a
plan's canonical entitlement block from the @hanzo/plans catalog (runs the
bundle 'entitlements' route; no data duplication, no fromLegacy re-impl).
- clients/world/entitlement.go: WorldLimits + WorldLimitsFromEntitlements
(pure) + ResolveWorldLimits(ctx, planID) — values sourced from world.*
entitlements, never hardcoded. FreeWorldLimits is the fail-closed floor
(catalog outage degrades to Free, never grants model/stream).
- GET /v1/world/limits?plan=<id>: machine-readable contract echo so agents/
dashboard self-config against the live catalog instead of hardcoding tiers.
- Tests: contract mapping (all tiers), fail-closed on unmounted catalog, and
end-to-end Entitlements against the real embedded bundle (world.model_api
present on pro/enterprise, absent on free).
Per-request enforcement (org->plan resolution + rate limiter wiring) is the
documented follow-up owned with feat/world-model-engine; both gates resolve
through ResolveWorldLimits so policy stays single-sourced.
Proxy capture of cloud->ring proved the Safe flow hits the ring's commit-after-
response read-after-write race TWICE, not once: createVault->createWallet ('vault
not found', already retried) AND createWallet->deploy ('wallet not found', which
502'd custody=safe). With ALL requests pinned to one node (via a debug proxy) the
deploy STILL 404'd, so it is a Postgres commit-visibility lag, not node affinity.
Extract doRetryNotFound(...notFound) (bounded 6x/250ms linear, ctx-aware, fail-fast
on any other error; do() only unmarshals on 2xx so out is safe across retries) and
use it for BOTH createWallet ('vault not found') and deploySafe ('wallet not
found'). go test ./clients/wallets/... green; cmd/cloud builds.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The embedded KMS write path POST /v1/kms/orgs/{org}/secrets defaulted a
missing env to "default" (envOr), committing the write to a bucket that
project/env/path readers (the kms-operator, cluster syncs) never resolve.
That split is what let an IAM z-password land in env=default while prod kept
serving the stale value. env is a first-class component of the storage key
(kms/secrets/{path}/{env}/{name}) and cannot be aliased, so a write with no
env now fails loud (400). GET/DELETE/LIST keep the envOr compat default (a
read/delete can't plant a value another reader trusts; legacy readers that
omit env must keep working). No PATCH route exists on this surface.
Regression tests: write without env -> 400 (and lands nowhere); write
env=prod is readable via the operator's project/env/path resolution (sha256
round-trip, values never printed) and is not visible in env=default. The
fail-closed-without-master-key test now sends a valid env so it still
exercises the 503 master-key gate rather than 400-ing on input.
Claude-Session: https://claude.ai/code/session_01D4FSvT3UfhrFJNQjrctjEj
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Enabled() staged path is now orthogonal to the Enable allowlist: a staged
subsystem (iam/ingress) mounts when named in EITHER Enable (strict allowlist) OR
the new EnableStaged (additive). CLOUD_ENABLE_STAGED=iam + empty CLOUD_ENABLE =
all-non-staged prod default PLUS iam — the faithful iam-fold canary/cutover shape
with NO hand-enumerated allowlist that silently drops a newly-added subsystem.
Proven: TestEnabled_StagedActivatesAdditively (iam on, non-staged default intact,
unnamed staged sibling stays off).
Completes the #71 auth repair as a clean dep bump (the fix lives in ai + iam, not
the cloud tree):
- github.com/hanzoai/ai v1.802.0 -> v1.802.1-0.20260708185316-0321c35877f0
(ai#79 0321c358: self-heal get-account identity — stop degrading real logins to
u-<hash> guests; fail-closed 401). Pseudo-version pins the commit while the
semantic-release patch tag mints (1 commit ahead of v1.802.0).
- github.com/hanzoai/iam v1.31.18 already pinned in main (iam#109 fail-closed
guest-mint gate) — MVS keeps it over ai's older iam pin.
go mod tidy added the authentic gopsutil/v4 transitive hashes (iam util); go mod
verify OK; -mod=readonly CGO_ENABLED=0 go build ./cmd/cloud green.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Both the iam and ai casdoor-derived forks resolve their SQLite handle from the
SAME env key (dataSourceName) + the one beego web.AppConfig global. A deployment
sets dataSourceName for ai; with iam enabled, IAM's bootstrap resolved that same
value and xorm-opened ai's DB (auto-migrating casdoor tables into it) -> the
documented boot crash that pinned every post-embed release and kept iam staged.
IAM's conf already honors an IAM-scoped IAM_DATABASE_URL above the shared
dataSourceName; pin it to IAM's own iam.db under DataDir so the two forks get
independent stores, order-independent, NO fork edit. Operator override respected.
Unit-proven: TestIsolateDatabase (iam-owned DSN, != ai dataSourceName, respects override).
Makes console.hanzo.ai (go:embed console) authenticate its money surfaces: the
first-party IAM session cookie → validated principal (sessionAccessToken → v.validate),
RED-hardened (H2 Secure cookie, H3 same-origin bridge gate), on iam v1.31.18 (H1
session-regeneration + iam-main security fixes). Pairs with console v8.4.122 which
addresses billing/commerce/keys at the canonical bare /v1 in embed mode.
v1.31.18 is iam main (guard-leak stamp #108, guest-signin fail-closed #109, capauth
PermAttenuate #106) UNION the InitEmbed line UNION RED H1 (SessionRegenerateID on the
sign-in transition). The prior embed pin v1.31.17 had diverged off an old base and was
MISSING those iam-main security fixes, so pinning v1.31.17+H1 would have shipped the
money embed without them. v1.31.18 ships H1 + guard-leak/guest/capauth + InitEmbed
atomically into this binary's in-process IAM. Transitive indirect bumps (purego/
plan9stats/locafero/gopsutil-v4/viper/tidwall-match) are MVS-driven by iam v1.31.18.
H2 (HIGH) — pin the IAM session cookie Secure. clients/iamsvc/iamsvc.go derived
Secure from web.BConfig.Listen.EnableHTTPS, which is FALSE (the binary listens plain
:8000 behind the TLS-terminating ingress) → the session cookie shipped non-Secure.
The embed bridge turns that opaque sid into a money bearer (hk- mint, balance/top-up),
so a non-Secure cookie is capturable off any plaintext leg and replayable. Pinned
Secure: true (the deployed edge is always HTTPS).
H3 (MED-HIGH) — gate the ambient-cookie bridge to same-origin. billing.go/commerce.go
forward GET verbatim to commerce; a SameSite=Lax cookie still rides a top-level GET, so
a cross-site link could drive the victim's own money action if any commerce GET mutates.
validatedPrincipal now fires the session bridge ONLY for a same-origin request
(sessionBridgeSameOrigin: Sec-Fetch-Site same-origin|none, else Origin/Referer
host==Host) — refusing cross-site AND sibling-subdomain (same-site). Bearer/JWT-cookie
paths (non-ambient) are unaffected. +TestSessionBridgeSameOrigin (7 cases) green.
REMAINING for money: H1 (session-fixation — SessionRegenerateID on the IAM sign-in
transition) lands in hanzoai/iam (compiled into this binary); coordinating.
The go:embed console (console.hanzo.ai → cloud:8000) authenticates against the
in-process IAM, which sets an OPAQUE, httpOnly session cookie (cloud_session_id)
and stores the user's IAM-minted access-token JWT SERVER-SIDE against that session.
The console's Next BFF token-minting routes are stripped by the static export, so a
browser request to a cloud-native route (/v1/console/keys, /v1/billing/*) carries
only the session cookie — no bearer — and validatedPrincipal refused it, 401ing
every authenticated surface (API keys, billing, every product page = shell).
validatedPrincipal now resolves that session cookie to the server-stored access
token (sessionAccessToken via web.GlobalSessions) as a LAST RESORT (after Bearer/
Basic/JWT-cookie), then validates it through the SAME v.validate (sig/iss/aud/exp).
Identity is bound to the VALIDATED session: the client holds only an unguessable,
httpOnly sid; the session never asserts identity itself. No-op on gateway-fronted
binaries (a bearer is present) and on binaries with no IAM session manager
(web.GlobalSessions == nil) — tested. CSRF: cloud_session_id is SameSite=Lax, so a
cross-site request never carries it; and cookieTokenNames already establishes cloud's
JWT-cookie auth posture. This is the v8.4.5-flagged 'set the cookie the sanitizer
looks for' path, done cloud-side from the session store (no cross-repo IAM release).
RED review requested before it fronts money (session-fixation / CSRF surface).
topupConfig defaulted to the placeholder 36900; align to the
genesis-canonical Hanzo mainnet chain id 36963 (lux/genesis, and the
rest of cloud clients/treasury+wallets already use 36963). Still
env-overridable via HANZO_CHAIN_ID.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The observability surface was mounted THREE ways over the same /v1/o11y/* paths:
clients/observe (order 44), clients/o11y's o11yscope (order 69), and the
hanzoai/o11y wildcard runtime (70/71). observe also served /v1/settings/:product,
which is console product config, not observability. Collapse to one and one way.
- ONE owner of /v1/o11y/{logs,metrics,status}: clients/o11y's o11yscope (order 69).
observe's richer logic is folded IN so nothing is lost — the REAL per-org RED
metrics + LLM usage (metricsread.go, was a stub in o11y) and the two-view logs
(admin infra stdout / per-org request-from-traces). Tenant isolation preserved:
org is principal.Tenant bound as a positional ClickHouse param, the product is
shape-validated → alias-mapped (console slug → workload) → allowlisted
(knownServices, SSRF/injection boundary). observe's productAlias merged into
resolveService so no product loses data. Admin god-view gates on c.IsAdmin()
(== owner=="admin" SuperAdmin after SanitizeIdentity), never a per-org isAdmin.
- /v1/settings/:product moved OUT of observe into clients/settings (it is NOT
observability). Behavior/contract preserved verbatim from observe: {config,
secretKeys} shape, KMS ref orgs/{org}/settings/{product}/{key}, (org,product)
store isolation, secrets-to-KMS-or-fail-closed. Replaces the orphaned, divergent
clients/settings stub with the live behavior and wires it in (order 138).
- /v1/query does not exist (no registrant, no consumer) — nothing to fold.
/v1/observe/health was the auto-derived GET /v1/<id>/health for id "observe";
it vanishes with the subsystem (o11yscope gets /v1/o11yscope/health; the runtime
serves its own gate-exempt /v1/o11y/api/v*/health*). Both documented.
- DELETE clients/observe; drop its import; add clients/settings; fix the stale
subsystems.go o11y comment ("reverse proxy to the dedicated o11y Deployment" →
the embedded reality: scoped reads 69 + in-process runtime 71 + OTLP ingest 72).
Net -1037 LoC. cmd/cloud + cmd/hanzo build; clients/o11y + clients/settings tests
pass (20/20), covering tenant isolation, secrets-never-plaintext, product
validation, alias resolution, and route precedence over the wildcard proxy.
One coherent change to the subsystem-registration layer. Concrete types over
`any` at the call sites, indirection deleted, generics only where they remove
real duplication.
CHANGE 1 — kill the per-subsystem `any`-unwrap boilerplate
Every in-repo subsystem's init() hand-wrote the identical
func(app any, deps cloud.Deps) error { a, ok := app.(*zip.App); if !ok {…}; return Mount(a, deps) }
Add ONE adapter, cloud.Typed(func(*zip.App, Deps) error) MountFunc, that does
the *zip.App recovery in a single place (fail-closed, never panics). All ~50
subsystems collapse to `cloud.Register("x", n, cloud.Typed(Mount))` /
`cloud.RegisterWithShutdown(..., cloud.Typed(Mount), Shutdown)`. Redundant
shutdown wrappers dropped where Shutdown already matches ShutdownFunc; kept
only where a no-arg Shutdown() needs signature adaptation.
MountFunc's param STAYS `any` on purpose: the pinned external subsystem modules
(hanzoai/ai, authz, base, commerce, metrics, o11y, licensing) register with
`func(any,…)`, and a `func(any,…)` literal is not assignable to a
`func(*zip.App,…)` parameter — retyping MountFunc would break those modules at
compile time. The assertion is now central, not per-subsystem. MountAll takes
the concrete *zip.App (threaded from Serve).
CHANGE 2 — OwnsHealth flag replaces the "<name>svc" health kludge
Some subsystems serve their OWN fail-closed /v1/<name>/health; the generic
always-ok liveness route in Serve would shadow it. The old fix encoded routing
policy in the id ("kmssvc" parked the generic route at an unrouted path). Now
Register/RegisterWithShutdown take `opts ...Option`; cloud.HealthOwner sets
MountSpec.OwnsHealth, and Serve's generic-health loop skips a HealthOwner. The
id is once again the clean route name. Invariant now uniform and checkable:
a subsystem serves /v1/<name>/health IFF it registers cloud.HealthOwner.
Migrated every health-owner to it: kms, paas, s3 (named in scope) plus
analytics, console, platform, ml (same kludge) and notify, plans, pricing,
security (had coincidental id==route; security's real probe reports a rule
count the generic route was silently dropping). pickKMSClient gate + all
tests + stale comments updated from the "kmssvc"/"s3svc"/… ids to kms/s3/….
CHANGE 3 — clean package renames (no collision)
clients/paassvc → clients/paas, clients/projectsvc → clients/projects
(package decls, filenames, the sole importer, error strings, userAgent, and
doc refs repo-wide). clients/kmssvc + clients/tasksvc KEEP their package names
— the `svc` disambiguates the subsystem from the same-named library it imports
(clients/kms, hanzoai/tasks); their ids are already clean (kms via CHANGE 2,
tasks).
CHANGE 4 — generic pick[T]
The five identical co-resident-or-RPC-or-disabled resolvers (IAM, Base,
Commerce, O11y, MQ) collapse into one
pick[T](cfg, log, name, label, zapAddr, rpc func(string)T, disabled func()T) T.
KMS/AI/VFS/Payments/Vault keep bespoke pickers — their construction genuinely
differs (embedded store / gateway preference / S3-admin backend / never
co-resident), so they are left alone.
Verified: CGO_ENABLED=0 go build ./cmd/hanzo/ and ./cmd/cloud/ both exit 0;
go vet clean on every changed package; `hanzo --help` lists kms/paas/projects/
s3/tasks svc-free; cloud root + renamed + health-owner package tests pass; new
build_registration_test.go covers Typed + HealthOwner. Net −199 lines.
Picks up the increment-2 crypto hygiene: the PartyID<=ValidatorSetSize DoS
bound on the quasar/pulsar Finalize path (Item7a) + the structural-Verify lock
(Item7b). v1.35.32 corrects a 1-based off-by-one in v1.35.31 that rejected the
Nth validator; verified the controlplane N=7 ceremony finalizes under -race.
LOW severity (ingestLeg bounds PartyID upstream) but the fix is now live-pinned.
Add clients/ingress — an embedded edge plane in the cloud binary so the ONE
binary can BE the fleet edge: terminate TLS, run ACME, and reverse-proxy by Host
to upstreams, configured LIVE over /v1/ingress with no static routes.yaml and no
restart to change a route (hot-apply via an atomic engine snapshot swap).
Control plane (zip): /v1/ingress/{routes,services,middlewares,tls,status},
SuperAdmin-gated, per-tenant SQLite persistence, route Host globally unique;
every mutation reloads the engine.
Data plane (net/http): :80 (ACME HTTP-01 + router) and :443 (SNI TLS termination
via x/crypto/acme/autocert + router). Started only in edge role
(CLOUD_INGRESS_EDGE_ENABLED); app role keeps the listeners off — role = runtime
config, one binary.
Proxy: github.com/vulcand/oxy/v2 (Traefik lineage) weighted round-robin, the
Traefik router->service->middleware model. Middlewares: redirectScheme,
stripPrefix, addPrefix, headers.
STAGED subsystem (config.stagedSubsystems): linked but mounts ONLY when named in
CLOUD_ENABLE, so prod is untouched. Orthogonal to /v1/gateway (auth/rate-limit).
Build: CGO_ENABLED=0 go build ./... green; go test ./clients/ingress green (11 tests).
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Red-team findings on the write-fence primitive:
- HIGH: mirrorKey(plugin,shard) = "writefence/"+plugin+"/"+shard was not
injective — mirrorKey("kms/tenant-a","secrets") collided with
mirrorKey("kms","tenant-a/secrets"), so a push framed as one shard could
overwrite another's epoch/writer/payload. Real shard scopes carry '/'
(vfs replica.DBPath yields "projects/site"), so it is reachable. Fixed with
a %d:-length-prefixed key; TestMirrorKeyInjective_NoCrossShardAliasing locks
it (was red's failing PoC, now green).
- LOW: MinioConditionalStore.Get did StatObject+GetObject (two round trips);
tightened to one GET whose obj.Stat() ETag is consistent with the read bytes
— closes the window rather than relying on the CAS to absorb a stale version.
Core CAS/epoch soundness unchanged (red GO: N=64 same-epoch race → one winner,
retry-bounded, strict-> rejects epoch==recorded). Still shadow-only, unwired.
go test -race -count=200 green.
Red-pass finding: the check-#1 grep char class [A-Za-z0-9_, ] missed the
build-constraint negation form `-tags '!x,controlplane'` (the `!` broke the
match before reaching controlplane). Add `!` to the class so it is caught.
Verified the deeper guarantees hold (evasion-agnostic), so this is belt-only:
- ZERO non-test importers of clients/controlplane (grep-confirmed).
- The package has ZERO untagged files, so importing it into serve code fails
the untagged `go build ./...` — check #3 catches ANY -tags syntax, incl.
GOFLAGS=-tags=controlplane (verified: the pkg becomes buildable => check #3's
'matched no packages' assertion fails => CI red).
Runtime asserts + external-cert selfComposedCert seam confirmed wired through
guarded constructors. No stub-crypto path reaches a serve binary.
Self-review finding: containment.go's runtime guard trusts testing.Testing(),
which is backed by a linker-set string var (testing.testBinary, set by `go
test` itself per cmd/go/internal/load/test.go). Confirmed locally that
`go build/run -ldflags="-X testing.testBinary=1"` spoofs it to true in a REAL
(non go-test) binary — verified with a throwaway program before writing this.
containment.yml's grep step now also fails the build on any reference to
`testing.testBinary` outside the Go toolchain itself, so a build path that
tried to ship that spoof gets caught the same way a `-tags controlplane`
build path does. Documented as a known residual in the workflow's header:
this is a mitigation (CI catches it), not a cryptographic close — that needs
increment-2's real signing, tracked in doc.go.
Also fixed the exclusion patterns to be grep-implementation-agnostic (some
recursive greps don't prefix paths with "./"), verified against a planted
violation for both checks.
Closes the same-epoch double-write on the HIP-0107 data-plane push path
(github.com/hanzoai/vfs/replica, wired in internal/org): today the only
admission checks are replica.IsOwner (a pure local computation over a
possibly-stale membership view) and the StatefulSet Recreate deployment
shape (role.Role) — both comment-only, non-atomic, and the underlying
Store/Backend.Put is an unconditional overwrite ("Overwriting is allowed").
A deposed/partitioned writer and a freshly-elected one can both push.
internal/writefence/fence.go adds Fence.Push: a single atomic
read-check-CAS that (1) rejects any candidateEpoch <= the epoch currently
recorded for the shard (strict >, closing the same-epoch case) and (2)
performs the epoch-advance and payload append in ONE conditional write
against the store's live version token, so two racing writers cannot both
land — the store is the sole arbiter, never an in-memory cache. Retries
once on a lost CAS race, re-checking strict monotonicity against the new
state, so a same-epoch racer's retry fails ErrStaleEpoch rather than
silently duplicating the admit.
EpochSource is the pluggable seam clients/controlplane's lease epoch drops
into once it graduates from shadow (Stage 1 today) — this package imports
nothing from controlplane. ConditionalStore models the S3 If-Match / GCS
generation-match primitive; store.go backs it for real with minio-go's
native SetMatchETag/SetMatchETagExcept (already vendored at v7.0.100, no
go.mod bump). fake_test.go models the same semantics in-process with a
barrier hook that deterministically reproduces the concurrent-CAS race.
Tests prove: strict-epoch rejection of a same-epoch retry (same and
different writer), the raw CAS rejecting a race loser, the full
concurrent-Push race resolving to exactly one winner, a legitimately
higher epoch being admitted, a stale lower epoch being rejected, and
per-shard scoping. Not yet wired into the live push path (that remains
gated by controlplane's shadow flag per HIP-0116); this is the fence
primitive plus a precise wiring recommendation for hanzoai/vfs's block
layer, which currently exposes no conditional-write capability to adopt.
Stage-1 ceremony's crypto is stub/forgeable by design (doc.go); this closes
the drift risks doc.go's increment-2 worklist flagged:
- .github/workflows/containment.yml (PR-gated): greps every build/release
surface in the repo for `-tags controlplane` and fails the build if found,
plus a positive proof that `go build ./...` links clients/controlplane into
no cmd/ main and that the package still matches zero packages with no tag.
- containment.go: mustHarnessOnly fail-closed panics the moment this
package's stub crypto is touched (package-import-time for the
PartialZVerifier registration, construction-time for NewSigner/
NewStubComposer) unless ProductionBCCSigningReady() (hardcoded false) or
testing.Testing() (the Go toolchain's own go-test signal, unspoofable by a
real build) holds. Proven end-to-end via a real subprocess
(TestContainment_NonHarnessProcessRefuses), not just in-process logic.
- selfComposedCert typed seam (driver.go/signer.go): CertComposer.Compose now
returns an unexported wrapper only it can produce; verifyOwnCertStructure
accepts only that type, never a bare *quasar.QuasarCert. An externally-
received cert has no way to become one, so it cannot reach the structural
check even by mistake. VerifyExternalCert is the sole seam for such a cert
and fails closed (increment-2 crypto not implemented). Locked from a
black-box vantage in external_cert_test.go.
Containment verified unchanged: `go build ./clients/controlplane/...` (no
tag) still matches zero packages; `go build ./...` still links no cmd/ main
to the package; full `-tags controlplane -race` suite green, no test weakened.
The cloud-embedded console (console.hanzo.ai + team) is built from hanzoai/console
build:embed. hanzoai/console now ships <HanzoAnalytics/> (env-gated on
NEXT_PUBLIC_ANALYTICS_WEBSITE_ID). Default it to the console.hanzo.ai property
(7dce54ee-41f6-4751-96bf-fe005067c7c7, public per-site) in the console build stage
so the one native analytics tag renders on the next cloud build. GA4/Pixel off.
The luxfi/mpc threshold signer returns a NON-canonical r|s: s is frequently in the
upper half (s > N/2). luxfi/geth's tx validation (ValidateSignatureValues,
homestead=true) REJECTS high-S signatures, so the anchor's MPC-signed self-tx
failed on submit with 'invalid sender' (live: POST /v1/admin/treasury/anchor ->
status error, note 'submit: send tx: invalid sender'). recoverableSig now
canonicalizes s to N-s when it exceeds N/2 before searching the recovery id, so
the 65-byte r|s|v it hands tx.WithSignature is EIP-2-valid and recovers to the
treasury MPC wallet. Tests: TestRecoverableSig_LowSNormalization (forced high-S ->
low-S, still recovers). go test ./clients/wallets/... green; cmd/cloud builds.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Adversarially verify blue's class-A fixes hold under op COMPOSITION inside a
single block (which the original red suite exercised only as separate blocks or
single ops). Six hostile compositions — bare reassign, release+reassign,
release+assign, membership-remove+reassign, remove+release+assign, assign-steal
— are each refused end-to-end through the N=7 ceremony, and the live writer is
unchanged across every voter with its lease mirror consistent. Plus: the
authorized proven-dead handoff stays single-valued under redundant reassigns,
and assign+release of a fresh resource leaves no orphan writer (mirror desync
would be a second authority). GO: the double-write class is fully closed.
The ring's :8081 commits a newly-created vault to its DB AFTER writing the
createVault 201 response, so cloud's back-to-back createVault->createWallet (fired
microseconds apart on one keep-alive connection) races the commit and read-misses
the just-created vault -> 404 'vault not found' -> custody=safe 502. A slower
client (curl, separate processes) never observes the gap, which is why manual
repro succeeded. Bounded retry (6x, linear 250ms backoff, ctx-aware) on exactly
that 404; every other error still fails fast. Idempotent per attempt (fresh body).
go build ./cmd/cloud green; go test ./clients/wallets/... green.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Wires the #162 BindAnchorSigner seam to a real quorum signer. New:
- wallets.TreasuryAnchorSigner(org,chain): resolves-or-provisions the org's stable
KindTreasury wallet on the ring (reserved account 'treasury' / wallet
'reserve-anchor', idempotent) and returns its EVM address + a sign closure.
- The closure produces an EVM-recoverable r‖s‖v signature: the ring returns a bare
r‖s (64B, no recovery id) but tx.WithSignature needs 65B, so recoverableSig finds
the v whose recovery yields the wallet address (fails closed otherwise).
- POST /v1/admin/treasury/bind-anchor (global-admin): calls TreasuryAnchorSigner +
BindAnchorSigner, so subsequent /v1/admin/treasury/anchor commits the ledger root
signed by the treasury MPC wallet, not the lone KMS key. Returns the bound address
(fund it for gas on the Hanzo L1).
Tests: TestRecoverableSig (both parities recover to the signer) + _NoMatch (fail
closed). go build ./cmd/cloud green; go test ./clients/wallets/... green.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The o11y-scope landing added clients/o11y/{scope,status}.go referencing a metrics-
query layer (vmClient, newVMClient, promLabel, metricsQuery, queryMetrics,
metricsResult, metricPoint, usageRollup, boundRangeMinutes) whose source file was
never committed → `go build ./cmd/cloud` failed (undefined symbols), taking the
whole deploy plane down (no new cloud image buildable from main). Restore the file
to the surface's own honest-empty contract: newVMClient reads O11Y_VM_URL and an
unset/unreachable VM degrades every query to an honest-empty series (never a
fabricated point); status.go's VM up-inventory works when VM is wired. queryMetrics
returns the honest-empty RED series until the VM query_range wiring lands. Full
`go build ./cmd/cloud` now links; go test ./clients/o11y passes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Brings the embedded ai subsystem up to v1.802.0:
- #76 opt-in auto-routing (virtual auto/zen-router model, X-Routed-Model)
- #77 per-org enable/disable (OrgSettings precedence)
- #78 admin-settable defaults (reserved "*" row, /v1/get-routing-defaults)
+ RoutingEvent collection (no prompt text) + /v1/export-routing-ledger
Edge contract unchanged; auto_routing_billing_test green against the new
module (ok github.com/hanzoai/cloud). Pre-existing clients/o11y compile
break on main is untouched (fix lands separately).
Kill the second automation surface. /v1/auto was a per-org reverse proxy
(clients/auto + clients/auto/proxy) to the standalone hanzoai/auto engine
(auto.hanzo.svc) — a duplicate of the native, in-process /v1/automations
Connectors+Automations engine (clients/automations, cloud.EmbeddedTasks,
706-piece catalogue). One engine, one surface: /v1/automations is the ONE
native automation engine. The external engine + its console link-out are
retired (console + universe in paired PRs).
- Remove the order-140 blank import of clients/auto from subsystems.
- Delete clients/auto/ (auto.go + proxy/).
No functional loss: /v1/automations already serves flows/versions/runs/
pieces/MCP natively. clients/kb keeps its own AUTO_UPSTREAM piece-runner
coupling (a separate, pre-existing bridge to a never-implemented engine
endpoint) — reported for a follow-up, not touched here.
go build ./... green, go vet green, go test ./clients/automations + root ok.
* refactor(automations): rename connector catalogue pieces -> connectors (HIP-0125)
The automations connector CATALOG surface drops the ActivePieces term "pieces" for the ONE Hanzo term "connectors":
- GET /v1/automations/pieces -> /v1/automations/connectors; /pieces kept as a
byte-identical back-compat alias (same handler) so live clients never break.
- Catalog{PieceCount,Pieces} -> {ConnectorCount,Connectors}; PieceMetadata/
PieceAuth/PieceAction/PieceTrigger -> Connector*; JSON tags pieceCount/pieces
-> connectorCount/connectors; embedded catalog.json + OpenAPI updated to match.
- Test proves the /pieces alias mirrors /connectors byte-for-byte.
Deliberately UNCHANGED (persisted @xyflow builder wire contract; renaming would
break live clients + stored flows): the flow-step protocol PieceName/pieceName,
PIECE/PIECE_TRIGGER, corePiece. Aligning those is a staged migration (HIP-0125).
* chore(automations,git,framework): scrub AI-slop placeholder comments (Rob Pike pass)
Comment-only, zero behavior change. Removes agent-note narration and future-work hedges, keeps the real WHY:
- automations.go: drop "a separate agent later OVERWRITES this file" narration; keep the Catalog-is-the-wire-contract invariant.
- framework/naming.go: "value for now" -> "value derived from now" (it reads the now arg, not a hedge).
- git/git.go: drop TODO(billing) + "in the MVP" hedge; state the git.usage meter fact.
- git/storage.go: drop TODO(vfs)/MVP/follow-up narration; keep the WHY osfs (not vfs) is used (vfs.FS does not implement go-billy).
Kept as real WHY/invariants (not slop): connector_core.go loopback-test SSRF guard, connector_slack.go httptest override, affiliates/store.go sentinel + PendingCents; types.go was already cleaned in the rename commit.
* docs(automations): point connector-rename references at HIP-0126 (0125 was taken)
---------
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
CTO decision: ONE automation engine = the native Go /v1/automations
(clients/automations on cloud.EmbeddedTasks). This removes the redundant
/v1/auto reverse-proxy subsystem (clients/auto), a per-org proxy to the
standalone ActivePieces Deployment (auto.hanzo.svc).
- delete clients/auto/ (auto.go + proxy/)
- drop the order-140 blank import from subsystems.go
Safe: no live caller of cloud/v1/auto — console link-outs to auto.hanzo.ai,
and clients/kb calls the engine directly via its own AUTO_UPSTREAM client
(untouched here). The native /v1/automations surface is unaffected.
NOTE (does NOT retire the ActivePieces Deployment): clients/kb/sync_piece.go
still executes connector pieces via the engine at /v1/auto/pieces/{piece}/run;
the native engine exposes the piece CATALOGUE but not piece EXECUTION yet, so
auto.hanzo.svc must stay until native reaches piece-run parity.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Comment-only tightenings, zero behavior change:
- pubsub/o11y: drop the misleading "GC won't collect" narration on the
package-level server/collector refs; state the real reason (shutdown
reachability) or the actual invariant (metrics ref is a write-only keepalive).
- iamsvc: condense the 11-line InitEmbed block that verbatim-restated the
package doc down to the fail-closed WHY that matters at the call site.
No code changed (git diff: comments only).
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Safe deploy (POST /v1/wallets/{id}/smart-wallet) resolves the owner wallet by its
db.Wallet PRIMARY KEY (orm.Get). The :9800 internal /keygen mints a threshold key
but persists NO db.Wallet row, so deploy 404'd 'wallet not found' (live: every
custody=safe create -> 502). The ring's Safe surface is VAULT-scoped: the only
create path that persists a db.Wallet AND returns its id is
POST /v1/vaults/{id}/wallets.
safeCustody.Provision now: createVault -> createWallet (vault-scoped, returns db
id + internal WalletID + EOA) -> deploySafe(dbId). KeyRef stays
<internalWalletId>|<smartWalletId> (owner-sign via :9800 uses the internal id;
propose via :8081 uses the smart-wallet id); the db id is only needed for the
one-time deploy. safeclient gains createVault + createWallet; the stub test now
emulates the vault/wallet-create routes. go test ./clients/wallets/... green.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Add two org-scoped cloud-api surfaces for the enterprise console:
- /v1/usage/summary (clients/usage): the org's unified footprint roll-up —
spend by category over time + wallet (from the commerce ledger) plus LLM
usage totals (from the warehouse). Composes existing sources server-side;
each degrades independently to honest zeros with a source marker. Org from
the validated bearer only (principal.Tenant); a forged X-Org-Id with no
principal 401s and never reaches commerce.
- /v1/audit (clients/auditlog): the per-org twin of the admin god-view — an
org admin reads ONLY their own org's events off the SAME tamper-evident,
hash-chained store. Org PINNED server-side (a client ?org is ignored);
filters time/actor/action/resource/resourceId/result + pagination.
- audit: extract the shared audit.Wire projection (used by both the admin and
org routes, one JSON contract) and add a ResourceID filter to audit.Query.
Tests: usage (pure roll-up/categorization + HTTP scoping/honest-zeros),
auditlog (real in-memory recorder: scope isolation, filters, pagination,
401/501), audit (ToWire + ResourceID). CGO_ENABLED=0 go build + go test green.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Reviewing my own displacement fix adversarially: red tested release+reassign,
but a standalone release of a LIVE holder was still admitted, and after it the
shard is unowned — so release(victim) at H then assign(attacker) at H+1 puts a
second live writer on the shard (same class, pure policy, survives real crypto).
Fix: a LIVE holder's lease is immutable — releasable only when the holder is
proven dead (ErrUnauthorizedRelease), symmetric with reassign. Closes the whole
double-write class, not just red's two tested paths. + TestPolicy_ReleaseLiveHolderRefused;
TestRSM_DeterministicConvergence now marks the holder proven-dead (out-of-band)
before releasing. Suite green.
* treasury: anchor signs through a quorum-gateable seam, not a lone key
anchor_evm.go held the signer's private key in-process and did types.SignTx.
Decouple WHERE the key lives from the tx builder via a txSigner seam:
- keySigner — the existing local KMS-provisioned key (default; unchanged result,
proven byte-identical to types.SignTx).
- mpcSigner — delegates the 32-byte EVM signing hash to a quorum-gated custody
backend (the reserve's 3-of-5 treasury MPC wallet), bound via BindAnchorSigner
(the finance seam). The bound signer wins over any local key.
submit() now hashes the tx, delegates the hash to the resolved signer, and
applies the recoverable signature — agnostic to single-sig vs threshold. Fails
closed when neither signer is available (never fabricates a signature).
Test proves both paths recover to the correct sender and the quorum signer is
invoked exactly once; the live ring is a config swap.
* feat(gpu): BYO-GPU worker uploads render outputs to the org gallery
After studio.render completes on the local GPU, the worker fetches each finished
output from the local studio (/view) and POSTs it to the org studio's /upload/output
with the user's IAM bearer — landing it in orgs/{org}/output (S3-mirrored to the
gallery). No S3/rclone credentials ever touch the box; the session token is the only
credential. Upload target resolves from input.uploadUrl, then HANZO_STUDIO_UPLOAD_URL,
then studio.hanzo.ai. Proven end-to-end against studio 0.14.9 (aud hanzo-console).
* feat(gpu): per-machine share policy — advertised on the fleet record, enforced at claim
A linked GPU can be shared to specific orgs/projects/job-types/models with limits via
ONE policy object on the machine record (SharePolicy). It rides in the fleet
registration (input.policy) and is enforced ONCE, at claim: a job outside the policy
is failed back so an eligible worker takes it. nil/zero policy = fully permissive
(unchanged behaviour). Loaded from HANZO_GPU_POLICY (inline JSON) or
HANZO_GPU_POLICY_FILE. Unit-tested (reject matrix + loader).
Server-side multi-org queue fanout + metering-to-org+project remain follow-ups; the
worker enforces its own policy today (workers still claim their own org's queue).
* feat(world): GDELT + allowlisted-RSS news data plane (clients/world)
First vertical slice of the World news backend in the unified cloud binary:
GET /v1/world/news merged, filtered, freshest-first feed -> {items:[…]}
GET /v1/world/pipeline per-(org,project) pipeline config
PUT /v1/world/pipeline upsert feeds + keyword/region/source filters
GET /v1/world/stream SSE live refresh (ZAP-native, org+project scoped)
- Ports world/api/{gdelt-doc,rss-proxy}.js: GDELT 2.0 Doc artlist + host-
allowlisted RSS/Atom (~180-domain SSRF allowlist, enforced at PUT boundary,
at fetch time, and on redirect targets).
- Org/project isolation on every path (principal.Tenant/Project); SQLite
pipelines table PK(org,project); in-memory TTL feed cache (10m).
- RegisterWithShutdown order 142; one blank-import line in subsystems.go.
- Tests: httptest-stubbed upstreams (deterministic/offline) + live-verified.
---------
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Follow-up rails from the red pass + the cryptographer audit (all on top of the
class-A fixes):
- checkQuorumSafety(N,quorum,f) asserts N>=3f+1, quorum>=2f+1, 2q>N+f at cluster
construction (fail closed). The 2q>N+f margin at N=7 is exactly 1 and is the
whole basis of the no-fork property, so a future sizing change can never
silently break safety. + TestSafety_QuorumParametersAreByzantineSafe.
- verifyOwnCertStructure: renamed the driver's structural self-composed cert
check away from 'independent triple-gate verification' and documented that an
external cert must go through the cryptographic VerifyUnderPolicy (increment-2),
never this structural path (red #8).
- commitZ: domain-separate the z-share commitment by session + party
(H(cp-commit||sid||party||z)) so a commitment cannot be replayed across
sessions/parties (red #4 hardening).
- doc.go: record the red->blue outcome (no-fork core held; 4 class-A closed), the
CLASS-B caveat (stub secrets are public-seed-derivable -> safety suite meaningful
only under real crypto), and the increment-2 security worklist (distributed DKG,
authenticated handoff + KMS fence, RSM-level authz re-verify, external-cert
crypto verify, CI guard against -tags controlplane releases).
Suite green under -tags controlplane -race; default build unaffected.
Red found a CRITICAL double-write + 3 more class-A breaks (pure orchestration/
policy, survive real crypto) with failing exploit tests. All closed; red's 4
class-A tests now pass without weakening them; class-B (stub-crypto-forgeable)
deferred to the real-crypto increment with explicit t.Skip TODOs.
#1/#2 CRITICAL double-write (policy.go, placement.go): displacement of a LIVE
shard writer now requires proven-death by out-of-band evidence. A proposer-
written same-block release authorizes nothing (it is not holder consent), and
membership removal no longer manufactures proven-dead. Fail-closed increment-1
posture; authenticated graceful handoff + KMS fence are increment-2.
#3 HIGH barrier forgeable (driver.go, custody.go, signer.go, transport.go):
Round1 commitments are now proof-of-possession authenticated exactly as Round2
legs, so one node cannot forge a quorum of spoofed commitments to defeat
commit-before-reveal.
#4 MEDIUM apply fork gate (rsm.go): RSM.Apply re-checks ParentRoot == the
applied-state commitment, so a block that does not extend local state can never
mutate it (defense-in-depth for a future recovery/gossip path).
Corrected TestPolicy_ShardReassign_WithRelease (it asserted the vulnerable
same-block-release-authorizes-displacement behavior) to assert the fix. Updated
rsm_test blocks to extend state properly (the new parent-root gate). Suite green
under -tags controlplane; default build unaffected (package is tag-gated).
New KindSafe custody composes the ring's TWO planes without importing luxfi/mpc:
- :9800 internal threshold API (mpcclient) — keygen the owner MPC EOA + owner-sign
- :8081 product API (new safeclient) — CREATE2 Safe deploy + EIP-712 Safe-tx propose
safeclient.go mints a SHORT-LIVED HS256 ring JWT (iss=mpc.lux.network, aud=mpc-api,
role=admin, org-scoped) hand-rolled (crypto/hmac, no jwt dep) from the ring's
MPC_JWT_SECRET — resolved from cloud's in-process KMS via
CLOUD_WALLETS_MPC_JWT_SECRET_REF, NEVER a plaintext env value. The deploy route is
role-gated (owner|admin), so role=admin clears it.
safeCustody.Provision: keygen (owner EOA) -> deploy Safe(owners=[EOA], threshold=1)
on the wallet's EVM chain (per-wallet, default Hanzo L1 36963); KeyRef encodes both
ring handles (<mpcWalletId>|<smartWalletId>); Address = predicted Safe contract.
Sign: owner-approval signature via :9800 (uniform /v1/wallets/:id/sign). New route
POST /v1/wallets/:id/safe-tx composes the ring propose (EIP-712 MPC-sign) via a
safeProposer capability type-assert (no Kind switch). Fails closed
(ErrMPCNotConfigured) until CLOUD_WALLETS_MPC_API_ADDR + the JWT secret are wired.
Tests: TestSafeCustody drives a stub emulating both ring planes (asserts the minted
JWT is HS256-valid with correct iss/aud/role/org) + TestSafeCustody_FailClosed.
go test ./clients/wallets/... green; CGO_ENABLED=0 go build ./cmd/cloud green.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Two coupled changes so main builds green AND the MPC custody surface is live:
1. Fix the broken build on main. #160 (CLOUD_ROLE writer/reader HA split) and
#163 (Stage-0 control-plane inert config) each added a `Role` field to the
SAME Config struct on separate branches; the merge left Config.Role
redeclared (role.Role vs string) + a duplicate struct-literal key, so
`go build ./cmd/cloud` failed (release lane stuck at v1.786.124). Rename the
inert #163 field to ControlPlaneRole (env ROLE, consumed by nothing yet). The
HA Role (role.Role, CLOUD_ROLE, used by serve.go/build.go) is unchanged.
2. Register the wallets subsystem. clients/wallets (#151/#161) was never blank-
imported into subsystems.go, so its init() never ran and /v1/wallets was
unrouted (404) despite the code shipping. Add the order-127 blank import so
the accounts/wallets/custody/keys/sign surface mounts — KMS custody always
on; mpc/treasury fail closed until CLOUD_WALLETS_MPC_ADDR +
CLOUD_WALLETS_MPC_API_KEY_REF are wired. This is the seam the treasury anchor
(#162 BindAnchorSigner) binds through.
Verified: CGO_ENABLED=0 go build ./cmd/cloud green; wallets + config tests ok;
local boot logs 'wallets mounted' (defaultCustody=kms) then 'listening', no panic.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Two PRs landed on main that both added a Config.Role field — the #160
HA writer/reader role (role.Role, load-bearing in Serve) and the Stage-0
control-plane role (string). The text-merge compiled to a duplicate field
and broke the default build. Rename the inert control-plane field to
ControlPlaneRole (env ROLE unchanged); the HA Role keeps its name and all
cfg.Role.IsReader()/String() consumers are untouched.
* refactor(kms): de-alias badger→zapdb (the embedded store IS ZapDB)
clients/kms/kms.go imported the store as `badger "github.com/luxfi/zapdb"`.
The store is luxfi/zapdb — the canonical Lux embedded KV, a hardened Badger
fork whose Go package is still literally `package badger`. The alias made call
sites read like raw dgraph-io/badger. Rename the alias to `zapdb` so every call
site is self-documenting; behaviour is byte-identical (same package, same API).
Confirms the invariant: `grep -rn dgraph-io/badger` across cloud = 0. There is
no raw Badger anywhere; the one embedded store is ZapDB.
* feat(cloud): CLOUD_ROLE writer/reader split + read-only KMS reader + writer-pin
Introduces an explicit HA role so read replicas can be added WITHOUT ever
risking a second writer opening the RWO stores. Default is byte-identical to
today: unset CLOUD_ROLE ⇒ Writer ⇒ the single pod that owns the RWO PVC.
- role: CLOUD_ROLE ∈ {writer(default), reader}. Serve fails CLOSED on an
explicitly-invalid value (a wrong guess demotes the real writer or risks a
second one). Pure, tested, imports nothing from cloud.
- kms: Config.ReadOnly opens the ZapDB store READ-ONLY with the lock guard
BYPASSED — a reader serves secrets off a restored replica and NEVER takes the
exclusive write lock (the mechanism proven safe by luxfi/zapdb's
WithReadOnly + BypassLockGuard; zapdb-replicate uses the same to coexist with
a live writer). Reader with no restored store / no key fails closed. Tested
round-trip: writer writes → reader reopens read-only → reads back; reader
writes rejected.
- writerpin: the single-writer election seam. SingleWriter (production-correct
for StatefulSet replicas:1) is the default; ConsensusPin (Quasar leaderless
election) is an HONEST stub that fails closed with ErrNotImplemented rather
than fabricating a pin. Tested.
- wiring: Serve resolves+logs the role and the backing pin; pickKMSClient opens
KMS read-only for readers. Writer path unchanged.
NOT YET wired (reported for Red/CTO): consensus election (writerpin gates no
store-open yet — k8s guarantees the single writer); reader gating of the
audit chain / durable tasks / per-tenant SQLite (still open writable) — the KMS
reader path is the completed slice. Data replication runs as sidecars at the
manifest layer (hanzoai/replicate for SQLite, luxfi/zapdb-replicate for ZapDB),
not via in-process import.
* feat(ha): fail-closed reader write-guard + prove in-process KMS backup
ReaderGuard: one boundary middleware rejects mutating verbs on a Reader
(405), gating EVERY store (KMS+audit+tasks+SQLite), not just KMS's
read-only open — a mis-routed write can no longer silently persist to a
reader's ephemeral dir and vanish on restart (H4). No-op on a Writer.
replication_test: real *zapdb.DB writer streams incremental age-encrypted
db.Backup blocks WHILE live; reader Restores into its OWN separate dir —
refutes the C1 'second-process open fails' path and proves the producer.
Fail-closed test: no recipient => no block (never plaintext to S3).
* test(ha): reader-guard verb matrix + replication edge cases
ReaderGuard: GET/HEAD/OPTIONS reach the store, POST/PUT/PATCH/DELETE all
405 without reaching it; Writer path (guard unmounted) serves every verb.
replication: wrong-identity restore fails closed; restore requires manifest
+ identity (unhydrated store never serves empty); repeated/no-op/overwrite
backups restore to the exact latest value (chain-correctness invariant).
* test(config): align IAM single-replica test with staged-subsystem contract
The 'empty list -> iam-enabled' subtest predates IAM becoming a STAGED
subsystem (stagedSubsystems["iam"]=true): the empty-Enable mount-all
default deliberately does NOT mount IAM (it corrupts the shared Beego
global and crashes `ai` with SQLITE_CANTOPEN). So empty list is
iam-DISABLED and >1 replica is allowed; the guard fires only when iam is
EXPLICITLY enabled. Code was correct; the test asserted the pre-staging
behavior. Pre-existing red on main, unrelated to the HA change.
Promote the luxfi consensus stack (consensus v1.25.15, bft v0.1.5,
p2p v1.21.1, validators v1.2.0) from indirect to direct requires, and add
four INERT control-plane config fields. Zero behavior change, reversible.
Deps + inert config only — no engine imported/started, no routes, no
serve.go/build.go behavior change.
v1.25.15 is the minimal clean tag: it already carries NewBFT (consensus.go:168)
+ engine/bft, its graph pulls validators v1.2.0 (Manager), it requires exactly
pulsar v1.1.1 (which stays v1.1.1 — zero drift), and it is the MVS-selected
version, so promotion is a no-op to the compiled graph. A lower tag would
downgrade the whole build's consensus (behavior change); a higher tag drifts
pulsar + consensus code.
The four are held direct by controlplane_deps.go: blank imports behind the
never-set //go:build controlplane_deps tag, so nothing links into the binary.
go mod tidy keeps them direct (it reads all build tags); deleting the file
reverts them to indirect. NodeID/Peers/Role/ControlPlaneQuorum parse in
LoadConfig (NODE_ID/PEERS/ROLE/CONTROL_PLANE_QUORUM) but no subsystem reads them.
Architecture direction (proposed, not shipped): the control plane is designed
to run Quasar (post-quantum BFT, protocol/quasar Submit->Finalized) under a
strict-PQ cert profile with a Pulsar RoundSigner threshold signer.
tidy also corrected pre-existing drift on main (nats-io/nats.go indirect->direct
via clients/kafka/interop_test.go; pruned 7 superseded go.sum lines) — verified
identical on pristine origin/main.
Red review findings:
- go.sum: luxfi/precompile v0.5.37 zip hash disagreed with sum.golang.org
(h1:Yh3dJ+... vs authoritative h1:2v0z...) → cold-cache CI SECURITY ERROR.
Corrected to the sumdb-vouched hash.
- telemetry.go: remove the plaintext OTLP-HTTP fallback (newTraceExporter) that a
stray/standard OTEL_EXPORTER_OTLP_ENDPOINT could use to silently downgrade
tenant-carrying trace spans to cleartext. ONE wire now: ZAP. Dropped the
otlptracehttp import (also severs its transitive grpc pull) and the dead
otlpEndpoint parameter. OTLP stays only the collector's interop receiver.
The ai subsystem serves a virtual `auto`/`zen-router` model that resolves to a
concrete model id before pricing/billing, meters its own token cost keyed on the
SERVED model, and reports it via the X-Routed-Model header. The cloud edge prices
/v1/ai/* by PATH (0, self-metered), never by the request model, so `auto` bills
as whatever it resolved to — and the edge passes X-Routed-Model through untouched.
- auto_routing_billing_test.go: TestAutoRoutingBillsAsResolvedModel (edge does not
double-bill /v1/ai/* + header pass-through) and TestDefaultPriceAiPathModelAgnostic.
- AUTH_BILLING_CONTRACT.md §4a: document the binding.
No code change needed — cloud already meters ai from the subsystem's own usage
record (which keys off the resolved request.Model), so the edge binds correctly.
Route ExportTraceServiceRequest wire encoding through github.com/zap-proto/zap2pb
(the sanctioned ZAP<->protobuf boundary) instead of importing
google.golang.org/protobuf {proto,encoding/protowire} directly. Wire bytes are
byte-identical (repeated ResourceSpans under field 1); TestUploadTracesOverZAP
still decodes the spans over the real ZAP transport.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
anchor_evm.go held the signer's private key in-process and did types.SignTx.
Decouple WHERE the key lives from the tx builder via a txSigner seam:
- keySigner — the existing local KMS-provisioned key (default; unchanged result,
proven byte-identical to types.SignTx).
- mpcSigner — delegates the 32-byte EVM signing hash to a quorum-gated custody
backend (the reserve's 3-of-5 treasury MPC wallet), bound via BindAnchorSigner
(the finance seam). The bound signer wins over any local key.
submit() now hashes the tx, delegates the hash to the resolved signer, and
applies the recoverable signature — agnostic to single-sig vs threshold. Fails
closed when neither signer is available (never fabricates a signature).
Test proves both paths recover to the correct sender and the quorum signer is
invoked exactly once; the live ring is a config swap.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The prior mpcclient targeted a DECIDED-but-nonexistent dashboard route tree
(/v1/wallets/{id}/sign, /v1/treasury/*) authed with a hand-minted HS256 JWT.
The deployed luxfi/mpc ring's real, working server-to-server custody surface is
the internal threshold API (cmd/mpcd/main.go, :9800): POST /keygen + POST /sign,
gated on the static MPC_INTERNAL_API_KEY bearer token — the exact contract the
ring's own /sign handler documents for a custody adapter.
Reconcile cloud to that contract:
- mpcclient.go: keygen + sign over the internal API; static bearer key (KMS),
no JWT/dependency; deterministic idempotency key per (org,wallet,digest).
- custody.go: mpc + treasury provision via keygen, sign via /sign with the
wallet's EVM chain id; Rotate preserves the address (ring-managed shares).
Treasury quorum governance moves to the finance policy layer over this same
primitive (no separate ring route).
- wallets.go: CLOUD_WALLETS_MPC_API_KEY_REF (KMS ref of the bearer key).
- test: stub emulates the internal /keygen+/sign contract.
Feature-flagged: unset CLOUD_WALLETS_MPC_ADDR ⇒ mpc/treasury fail closed.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The console per-product Metrics dashboard groups usage on metadata.product /
metadata.agent, but commerce RecordUsage persists only provider/model (no product
field), so the breakdowns rendered honest-empty even though every non-LLM product
already meters+gates per-org via ResourceMeter (provider=<product>, default fee
$1.00, fail-closed 402 on zero balance).
clients/billing/usage.go is the ONE read-side adapter: usage() injects a canonical
metadata.product onto each ledger row (agent->agents, provisioning->kind,
token-metered->inference, else provider) from the SAME charged ledger, and honors
the previously-ignored ?product=<id> (server-side filter) and ?groupBy=product
(per-product spend rollup {product,requests,amountCents}). A row already carrying
metadata.product/agent wins, so it degrades to a no-op once the meter/commerce
persist them natively (forward-compatible).
No change to what is charged or gated; the balance floor stays enforced by default.
scopedBillingQuery is extracted so proxy() and usage() build the subject boundary
one way. AUTH_BILLING_CONTRACT.md documents coverage + the native-field checklist.
Tests: productOf table + enrich/filter/group units + handler-level ?product= /
?groupBy=product through the real route (33 billing tests green).
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The console image stage now FAILS the build when build:embed does not emit a
real static bundle (non-empty out/index.html + out/_next), instead of silently
degrading to the committed fallback shell. A broken console export can no longer
ship the placeholder to prod. Escape hatch: --build-arg ALLOW_PLACEHOLDER=1 for
a pure-Go dev image with no Node console.
hanzoai/console build:embed produces a real 7.7M static export (361KB index.html
+ 4.3M _next chunks); //go:embed bakes it into the ONE cloud binary. Also drops
the last console2 references (repo is hanzoai/console).
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The frontend repo is hanzoai/console (console2 was renamed away). Kill the
dead name across the build path + source so there is one name, one way:
- Dockerfile: clone hanzoai/console.git; ARG CONSOLE_REPO / CONSOLE_REF
- Makefile: CONSOLE_DIR; webui + build-standalone targets
- config.go: drop dead console2.hanzo.ai from the ZAP-WS origin allowlist
- comments across clients/* reference the console repo + its TS modules by
their real name
No behavior change beyond dropping one unused CORS origin. Root pkg builds.
Red review = SHIP; these close the 4 cloud-side low findings so the PR lands
with no known edges.
low-1 (rollback atomicity): createDedicated's inject-failure branch now calls
removeAddonURL BEFORE tearing the backend down. injectAddonURL is not atomic —
a strategic-merge PATCH can LAND server-side yet still return err (dropped
response / post-commit timeout); scrubbing the maybe-written <KIND>_URL first
means a committed-but-errored inject can't leave the instance pointing at a
deleted backend (a dangling DSN is worse than Base). Proven by
TestDedicated_InjectPartialWriteRollsBackOrphanKey (fake now models the
write-then-error partial failure; asserts inject THEN remove ran, key gone).
low-2 (kv fail-open corner): TestDedicatedKV_RequirepassEnforced boots the REAL
ghcr.io/hanzoai/kv image with the exact engine.args + mounted requirepass
config and asserts an UNAUTHENTICATED PING is REJECTED, then that default:<pw>
authenticates — locking down the one corner where, if the image ignored the
positional config, the instance would boot unauthenticated. Raw RESP over TCP
(zero new client deps); gated on CLOUD_KV_SMOKE_IMAGE + docker so the default
suite stays green, real in CI.
low-3 (strategic-merge sibling preservation): TestPatchAddonSecret_RealAPIServer
runs the ACTUAL k8sOrchestrator addon methods against a REAL kube-apiserver
(controller-runtime envtest) — inject KV_URL then SQL_URL => BOTH survive in
.data; RemoveAddonSecretKey drops one, keeps the other; idempotent on absent
key/Secret. Replaces the fake orchestrator's assumption with a server-proven
fact. Gated on KUBEBUILDER_ASSETS (skip without envtest binaries). Adds
controller-runtime v0.23.3 as a TEST-ONLY dep — pinned to the release that
keeps k8s.io at v0.35.3 (NO production client-go bump).
low-4 (datastore tag symmetry): dedicated datastore image tag floating ':26' ->
env("CLOUD_DEDICATED_DATASTORE_TAG", "26.2.3.2"), symmetric with sql/kv/docdb.
A floating ':26' resolves to whichever datastore lineage (bridge vs fork, distinct
data dirs) last pushed under it — a per-org instance must boot a deterministic
image.
go build ./... green; go test ./clients/provisioning/... green (envtest PASS
against a live apiserver, kv-smoke skips without docker).
(cherry picked from commit 04d841c4906b58fa06b1bc407b55c97e6661f169)
* feat(o11y): wire native datastore metrics ingest into the embedded runtime
Bumps hanzoai/o11y to the native datastore metrics driver and starts an
in-process ZAP metric receiver (clients/o11y/metrics.go) that writes metrics to
the datastore over upstream ch-go via o11y/pkg/datastoremetrics — no histogram
fork. Reuses the embedded runtime.TelemetryStore.ClickhouseDB() connection, so
the query plane (read) and metrics (write) share one datastore conn.
Opt-in + fail-soft: gated on O11Y_METRICS_ZAP_LISTEN, a no-op until set, errors
logged and swallowed so metrics ingest can never take the query plane down. This
unblocks retiring the standalone signoz-otel-collector metrics path once verified
(verify-then-cutover). CGO_ENABLED=0 build + vet + existing o11y/observe tests green.
* chore: re-pin o11y@main (native datastore metrics driver merged)
---------
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Red review = SHIP; these close the 4 cloud-side low findings so the PR lands
with no known edges.
low-1 (rollback atomicity): createDedicated's inject-failure branch now calls
removeAddonURL BEFORE tearing the backend down. injectAddonURL is not atomic —
a strategic-merge PATCH can LAND server-side yet still return err (dropped
response / post-commit timeout); scrubbing the maybe-written <KIND>_URL first
means a committed-but-errored inject can't leave the instance pointing at a
deleted backend (a dangling DSN is worse than Base). Proven by
TestDedicated_InjectPartialWriteRollsBackOrphanKey (fake now models the
write-then-error partial failure; asserts inject THEN remove ran, key gone).
low-2 (kv fail-open corner): TestDedicatedKV_RequirepassEnforced boots the REAL
ghcr.io/hanzoai/kv image with the exact engine.args + mounted requirepass
config and asserts an UNAUTHENTICATED PING is REJECTED, then that default:<pw>
authenticates — locking down the one corner where, if the image ignored the
positional config, the instance would boot unauthenticated. Raw RESP over TCP
(zero new client deps); gated on CLOUD_KV_SMOKE_IMAGE + docker so the default
suite stays green, real in CI.
low-3 (strategic-merge sibling preservation): TestPatchAddonSecret_RealAPIServer
runs the ACTUAL k8sOrchestrator addon methods against a REAL kube-apiserver
(controller-runtime envtest) — inject KV_URL then SQL_URL => BOTH survive in
.data; RemoveAddonSecretKey drops one, keeps the other; idempotent on absent
key/Secret. Replaces the fake orchestrator's assumption with a server-proven
fact. Gated on KUBEBUILDER_ASSETS (skip without envtest binaries). Adds
controller-runtime v0.23.3 as a TEST-ONLY dep — pinned to the release that
keeps k8s.io at v0.35.3 (NO production client-go bump).
low-4 (datastore tag symmetry): dedicated datastore image tag floating ':26' ->
env("CLOUD_DEDICATED_DATASTORE_TAG", "26.2.3.2"), symmetric with sql/kv/docdb.
A floating ':26' resolves to whichever datastore lineage (bridge vs fork, distinct
data dirs) last pushed under it — a per-org instance must boot a deterministic
image.
go build ./... green; go test ./clients/provisioning/... green (envtest PASS
against a live apiserver, kv-smoke skips without docker).
(cherry picked from commit 04d841c4906b58fa06b1bc407b55c97e6661f169)
Extend the dedicated-instance strategy so all four on-demand data add-ons —
Hanzo KV / SQL / DocDB / Datastore — route through ONE mechanism, and bind an
enabled add-on to an app instance by injecting its DSN as <KIND>_URL into the
instance's addons Secret (disabling reverts to Base).
- store: additive instance column (idempotent ALTER, threaded through Resource/
cols/scan/Insert) + ListByInstance(org,instance).
- dedicated: add sql (Datastore type=postgresql, POSTGRES_* env, PGDATA subdir)
and kv (type=valkey, per-instance requirepass via a MOUNTED config Secret since
the kv-server binary reads no password from env; DSN user=default). Engine
gains adminUser/env/args/secretMount so the CR builder stays one code path.
- addon_inject: injectAddonURL/removeAddonURL + orchestrator PatchAddonSecret
(strategic-merge, create-if-absent, key-preserving) / RemoveAddonSecretKey
(JSON-merge delete, idempotent). Reloader annotation + rev bump on the Secret.
- create: instance bind field (validated); inject AFTER the row Insert as part of
the atomic provision (rollback on failure). drop: revert to Base BEFORE tearing
the backend down.
- sql/kv move off the shared-logical registry (each org OWNS its instance); the
orphaned shared postgres/redis provisioners + pgx/go-redis direct deps removed.
Tests: instance column round-trip + ListByInstance isolation; sql/kv DSN + CR
shape; inject merges (second add-on never clobbers the first); un-bound create
skips injection; drop removes URL before teardown; inject failure rolls back the
whole provision. go build/vet/test green.
Pulls ai's additive isGlobalAdmin field on /get-account so console
recognizes global admins. Pure dependency bump: re-pins re-tagged
luxfi/* modules from source (GOPRIVATE, sumdb-bypassed) after the
documented content-hash drift, prunes cloud.google.com/go/compute and
stale hanzoai/iam v1.31.16 (ai dropped the GCP SDK and requires iam
v1.31.17). go build ./... green (CGO_ENABLED=0).
Fold the standalone otel-collector Deployment into the unified cloud binary:
an in-process OpenTelemetry Collector accepts OTLP (grpc :4317, http :4318) and
writes spans+logs into the same ClickHouse datastore cloud already reads for the
o11y query plane (signoz_traces / signoz_logs, cluster insights). Consumers point
at cloud.hanzo.svc instead of otel-collector.hanzo.svc.
Trimmed, driver-compatible pipeline (reuses the signoz clickhouse exporters that
compile against cloud upstream clickhouse-go v2.44.0):
otlp -> memory_limiter, resource(namespace=hanzo, env), batch
-> clickhousetraces (traces), clickhouselogsexporter (logs)
- OFF by default (CLOUD_OTLP_INGEST_ENABLED); fail-soft; ShutdownFunc flushes.
- DSN via env (envprovider), never on disk; metrics self-telemetry off so only
:4317/:4318 bind (no :9090 class clash).
- telemetry.go: add OTLP-HTTP exporter path so cloud can loop back to the
in-process ingest at localhost:4318 (ZAP stays default/canonical).
DEFERRED: metrics pipeline (signozclickhousemetrics) needs SigNoz dd-sketch
ch-go fork (chproto.DD/Store/IndexMapping) that will not compile against cloud
upstream ch-go; metrics ingest stays on the standalone collector. See
clients/o11y/LLM.md.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
One custody seam over three orthogonal signing backends selected per-wallet by
Kind:
- KindKMS single-sig custody IN-PROCESS via the embedded luxfi/kms client
(deps.KMS). The fully-exercised spine: a real secp256k1 key is generated, its
private bytes sealed under the KMS envelope, and every Sign recovers to the
wallet address. No network hop.
- KindMPC / KindTreasury custody DELEGATE over HTTP to the deployed luxfi/mpc
cluster via a thin typed REST client (the clients/mpcseal precedent). cloud
never imports github.com/luxfi/mpc. Unconfigured -> fail closed
(ErrMPCNotConfigured); a signature is never fabricated.
Config seam: KMS always available; mpc/treasury built only when
CLOUD_WALLETS_MPC_ADDR is set and the HS256 JWT secret resolves from a KMS ref
(never a plaintext env). Per-tenant SQLite (org column on every row, every query
filtered by org). Finance seam (WalletForLedgerAccount) is a pure lookup only.
Tests (incl -race): KMS single-sig end-to-end (sig recovers to address, sealed
at rest, rotate changes address), per-tenant isolation, custody seam selects
backend (fail-closed mpc/400 unknown), mpc path wired against a faithful stub.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Reconciles TestGlobalAdminGate_RequiresAdminOrgAndIsAdmin with the task #51 decision
to pin IAM_ADMIN_ORG to the operator org (hanzo). The assertions are unchanged — the
gate is owner==adminOrg AND isAdmin — but the comment/labels no longer editorialize
that owner==admin is the only valid adminOrg. adminOrg is deployment config; the test
pins it to "admin" hermetically and proves the two invariants that hold for ANY
adminOrg: isAdmin is required (a non-admin in the admin org gets nothing) and owner is
required (an admin of any OTHER org gets nothing). Renamed the different-org case off
"hanzo" (which prod now pins AS the admin org) to a neutral "globex" so it reads
unambiguously.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The finance ledger of record ran on a single process-wide {DataDir}/treasury.db.
Select the store per request from the validated IAM owner instead, so every
tenant's books live on their OWN Hanzo Base file and one tenant's writes can
never appear in another's read.
- sqlstore.Manager: opens+caches one *Store per tenant (mutex-guarded map). The
house/reserve ledger is one fixed file ({DataDir}/treasury.db, preserved — no
migration of live reserve capital); customer ledgers are {DataDir}/finance/{slug}.db.
- tenantSlug: injective (never folds acme/ACME), path-traversal-guarded, reserves
the house slug. Verbatim stem for a DNS-ish org, else a sha256 slug. Consumes the
treasury's canonical hanzoai/sqlite opener (Open) — ledgercore's per-tenant opener
is a test-only helper that would double-register the sqlite driver.
- treasury.Mount binds the ledger of record to the HOUSE store; myAccounts reads the
caller's OWN per-tenant file (house scope still honours the Formance/Postgres opt-in).
- StorageDriver(): one place decides the driver — sqlite (default, what prod runs) or
postgres (opt-in via FORMANCE_LEDGER_URL). Postgres option preserved, never the default.
Tests: per-tenant isolation (A's write never in B's read; distinct files; cache
identity; traversal stays in-dir; no case-fold) + default-driver=sqlite + opt-in
preserved. go build ./cmd/cloud green; ./clients/treasury/... green (incl -race).
Co-authored-by: hanzo-dev <dev@hanzo.ai>
* fix(config): stage IAM off the mount-all default — unblock the release boot smoke
Every cloud release since the IAM embed (#142) has failed its boot smoke and
the fleet stayed pinned to a pre-embed image (v1.786.110), so the treasury +
finance merges (#143/#144/#145/#147) never shipped.
Root cause (from the failed release smoke logs): with CLOUD_ENABLE unset the
binary mounts every registered subsystem, so iamsvc.Mount now runs
iamserver.InitEmbed(). In the smoke/Docker env InitEmbed panics opening its own
SQLite (IAM_DATA_DIR=/data/iam absent on the tmpfs) and is recovered to a
fail-closed 503 — but IAM and the ai subsystem are sibling casibase/casdoor
forks linked against the SAME beego module, so InitEmbed's half-initialised
shared process-global (web.BConfig / xorm adapter) then makes ai's own
bootstrap fail identically:
iam ERROR iamserver.InitEmbed: bootstrap panicked: unable to open database file (14)
ai INFO ai: initializing runtime
cloud: mount: mount ai: ai: bootstrap: unable to open database file (14) -> SMOKE FAIL
This would crash api.hanzo.ai in prod too (CLOUD_ENABLE is unset there), not
just the smoke.
Fix: make IAM a STAGED subsystem — excluded from the empty-Enable mount-all
default, mounted ONLY when named in CLOUD_ENABLE. This is exactly the HIP-0106
staged-rollout contract iamsvc already documents ('operator adds iam to
--enable only after the fold is verified'), now enforced in code. It restores
the pre-#142 mount-all set (iamsvc is the only subsystem #142 added to it), so
the boot smoke goes green again; hanzo.id keeps being served by the standalone
iam pod until an explicit, verified cutover. Local mount-all boot now reaches
'listening' with iam 'subsystem disabled' and ai mounted clean.
pickIAMClient already falls back to the remote/disabled IAM client when
Enabled("iam") is false (build.go), which is current prod behaviour, so no
deps.IAM regression. One activation mechanism (the enable-list), one place.
* test(identity): lock global-admin = owner==adminOrg AND isAdmin
The cloud admin surfaces (incl. /v1/admin/treasury/*) grant global admin only
to a validated principal whose org IS the admin org AND whose token carries
isAdmin. This locks that invariant end-to-end through the real JWKS-validated
SanitizeIdentity boundary, with the two cases that matter for the treasury flip:
- a hanzo-org ADMIN (owner=hanzo, isAdmin=true) -> NOT global admin
- a NON-admin in the admin org (owner=admin, isAdmin=false) -> NOT global admin
The sole global admin z@hanzo.ai is global admin because IAM promotes @hanzo.ai
into the admin org (owner==adminOrg), NOT because it lives in 'hanzo'. This test
is the guard proving the boundary must NOT be widened to owner==hanzo (e.g.
IAM_ADMIN_ORG=hanzo), which would elevate every hanzo-org admin to see all
tenants' finances. The gate stays owner==adminOrg AND isAdmin; the fix for z is
that its token carries owner=admin, never a wider gate.
---------
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The hanzo CLI (cmd/hanzo) had no build target, so a naive
`go build ./cmd/hanzo` used the machine default CGO_ENABLED=1 and panicked
at init: "sql: Register called twice for driver sqlite".
Root cause: with CGO on, github.com/hanzoai/sqlite (the canonical Hanzo
driver, imported by ~15 clients/*/store.go) compiles its mattn/SQLCipher
backend and registers "sqlite"; the embedded upstream deps that import
modernc.org/sqlite directly (base/core, o11y, commerce/db, orm/db) register
"sqlite" a second time -> panic.
Fix: build cmd/hanzo the same pure-Go way cmd/cloud and the Dockerfile
already ship (CGO_ENABLED=0). hanzoai/sqlite's !cgo backend IS modernc, so
the fork and every modernc importer resolve to a single registration. This
extends the existing CGO_ENABLED?=0 policy (see Makefile header) to the new
binary instead of adding a second way to build; no dependency is dropped and
hanzoai/sqlite stays canonical.
make hanzo # -> ./bin/hanzo, pure Go, one 'sqlite' registration
clients/treasury/anchor{,_evm}.go (Phase 2 ledger-root anchor) import
luxfi/geth and luxfi/crypto directly, so they are no longer indirect.
go mod tidy result; no version change.
Collapse the treasury's separately-written double-entry SQL onto ledgercore
(github.com/hanzo-fi/ledger) — the SAME engine the ledger's own store uses — so
there is exactly one double-entry implementation across the stack (church of
Rich Hickey: one double-entry value, not three places).
- Reimplement clients/treasury/ledger/sqlstore to back the ledger.Store/ledger.Tx
port with ledgercore instead of hand-rolled treasury_postings SQL. The
accounting truth — every balance and the reserve overdraw guard — is now
ledgercore's (postings -> moves -> balances + hash-chained log, idempotency-key
dedup, WithTx atomic read-then-write). The adapter only maps the treasury's
vocabulary (int64 cents, Kind/Program/Ref key, signed-Posting Entry) onto it.
- KEEP the port/adapter seam: Open()'s signature is unchanged, so treasury.go and
the Formance-HTTP opt-in are untouched — native (ledgercore) stays the default
backend. The engine (ledger.go) and the on-chain Root are UNCHANGED: each Entry
is round-tripped verbatim (as ledgercore transaction metadata), so the Root is
byte-identical to the previous store's, independent of ledgercore's own postings.
- Policy (revenue-share bps) stays in a small side table — it is Hanzo config, not
double-entry accounting, so it does not belong in the shared engine.
- Pin bun to v1.2.9 (replace): ledger-fi floors v1.2.18, which removed
schema.Formatter/NewFormatter/Append that hanzoai/o11y still uses; ledgercore's
compiled closure uses no v1.2.18-only API, so v1.2.9 satisfies both. hanzo-fi/ledger
is pinned to the PR-3 branch commit until it merges.
Tests (all green, incl. -race): overdraw guard, at-most-once payout, snapshot
reconcile, scope isolation, and Tx rollback all pass unchanged against the
ledgercore-backed store. The whole cloud module builds under -mod=readonly, and
the treasury test binary links NO modernc driver (so it does not reintroduce the
"sqlite registered twice" panic).
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The 36963 coreth fee market pins a 25 gwei min base fee, so a legacy tx priced
at base+1 strands the moment the base fee ticks up. anchor_evm.go now submits a
DynamicFeeTx (1 gwei tip floor, 2x-base-fee cap) — proven accepted on-chain as a
type-2 tx.
Adds clients/treasury/cmd/anchorctl: a one-shot in-cluster tool that provisions
the KMS-held signer (key -> KMS, only the address printed), funds it from a
genesis account, deploys contracts/TreasuryAnchor.sol, and can send anchor(bytes32).
Includes the compiled TreasuryAnchor.bin (solc 0.8.26, optimizer 200, cancun).
Deployed live: contract 0x53141dF42DF13Aad0512f2F08c3E3216EEFac5F2, owner = signer
0x703D4227d58d0b6A20BD721c940CED170470f634 (KMS ref hanzo/treasury-anchor/TREASURY_ANCHOR_SIGNER_KEY).
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The cloud registers the "sqlite" driver exactly once in every build mode EXCEPT
a naive `go test -race`: -race forces CGO=1, which links the fork's mattn
"sqlite" (github.com/hanzoai/sqlite) ALONGSIDE the embedded deps that import
modernc directly (ai/base/commerce/o11y/orm), so both register "sqlite" and the
binary panics at init ("sql: Register called twice for driver sqlite") — the
pre-existing failure in clients/{graph,kmssvc,o11y}.
`make test` (CGO=0) and `make test-cgo` (-tags sqlite_purego) already avoid this
by resolving the whole binary to modernc's single registration. This adds the
missing peer for the race detector: `make test-race` runs
`CGO_ENABLED=1 go test -race -tags sqlite_purego ./...` — CGO on for the race
instrumentation, but the fork forced to its pure-Go backend so mattn never
registers and "sqlite" is registered exactly once. The ONE way to race-test the
cloud.
Proof: `go test -race ./clients/o11y/` panics; `go test -race -tags sqlite_purego
./clients/{graph,kmssvc,o11y}/` all pass.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The finance.hanzo.ai + console Finance surfaces render real per-org data
instead of preview stubs. This adds no billing system — it PROJECTS the two
that already exist (the commerce customer wallet + the treasury reserve fund)
into the @hanzo/finance-ui contract (USD cents, optional-safe), scoped to the
validated IAM owner.
clients/billing/finance.go — six commerce-projected reads, reusing this
package's commerceProxy + per-org subject-pinning (one commerce read path):
GET /v1/finance/balance commerce balance (holds -> pendingCents)
GET /v1/finance/credits commerce deposit rows (grants, positive)
GET /v1/finance/usage?range= commerce withdraw rows -> series+lines+total
GET /v1/finance/invoices honest empty (no invoice ledger exists yet)
GET /v1/finance/payment-methods commerce portal, masked to brand+last4
GET /v1/finance/ledger?range= commerce ledger -> signed per-org postings
clients/treasury/treasury.go — GET /v1/finance/treasury reshaped from the
reserve Report into the TreasurySummary shape (reserve/committed/available +
honest Hanzo L1 anchor); the transparency policy rides along additively.
Tenant isolation: org A never sees org B (per-org subject pinned server-side,
client cannot widen scope); payment methods re-masked defensively so a PAN can
never leak. Honest empty/typed shapes where a data source does not exist yet.
Tests: go test -race ./clients/billing/... ./clients/treasury/... green;
go build ./cmd/cloud green.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
* feat(treasury): native double-entry reserve fund + backed-payout seam (#treasury)
The platform's OWN fund/reserve accounting, one layer ABOVE the per-org commerce
credit ledger. A store-agnostic, cloud-decoupled double-entry engine
(clients/treasury/ledger) — the SEED of the native hanzoai/finance central ledger
(the Go replacement for the Formance stack) — plus a Base/SQLite adapter
(ledger/sqlstore) and the cloud client (clients/treasury).
Core (clients/treasury/ledger): accounts + balanced journal entries (Σ postings==0,
refused otherwise), ONE shared fund:reserve pool with per-program payout sinks,
revenue-share policy (bps, one place), and the reserve GUARD — a fund debit that
would overdraw is refused, atomically, so growth-loop payouts are backed capital not
unbounded minting. Zero cloud/zip/SQLite imports; persistence is the Store/Tx port,
so it lifts to hanzoai/finance as a directory move.
Surface: GET /v1/treasury (org transparency), GET /v1/admin/treasury (report +
journal + anchor), POST /v1/admin/treasury/{policy,sweep,seed,anchor} (global-admin).
treasury.Reserve(program,ref,memo,cents) is the ONE seam the 3 loops call: backed →
proceed to credit; not backed → honestly pending; unmounted → passthrough
(backward-safe). Idempotent by ref (at-most-once fund debit). ledger.Root commits the
whole journal for the Hanzo L1 anchor (Phase 2 wires the KMS-signed submit).
Tests (-race, green): double-entry balances, revenue-share accrual + per-period
idempotency, reserve guard (backed→blocked), at-most-once, concurrent no-overdraw,
admin gate, Reserve passthrough+enforced, sqlstore round-trip + tx rollback.
* feat(finance): Formance ledger-of-record backend + backed payouts + scope-aware /v1/finance/*
Adopt Formance as the ledger of record behind a ledger.Backend PORT, without
reimplementing double-entry: two adapters satisfy the port — the native Base/SQLite
engine (offline/default, ships the reserve fund today) and clients/treasury/formance
(a real HTTP client to the Postgres-backed Formance Ledger v2 API: world→fund accrual,
fund→payout debit, 400 INSUFFICIENT_FUND→not-backed=the overdraw guard Formance
enforces, reference→idempotency). Select by FORMANCE_LEDGER_URL — a config flip. Root
computed via a SHARED hash so the L1 anchor is backend-agnostic.
Back the growth-loop payouts: referrals/affiliates/authors now DEBIT the reserve fund
via the ONE treasury.Reserve seam before crediting the recipient wallet — fund down,
wallet up, reconciled. Not backed → honestly pending (referrals) or 402 + VoidPayout
restores pending (affiliates/authors). Idempotent by ref (at-most-once). Unmounted →
passthrough (backward-safe; existing loop suites stay green).
Scope-aware /v1/finance/* — ONE engine, three tenancy surfaces (admin/console/finance
product): tenant derived from IAM, house/reserve locked to global-admin under
/v1/admin/finance/*, per-org callers see ONLY their own org:<tenant>:* accounts.
GET /v1/finance/accounts (per-org; admin ?scope=house|?org=<t>). Storage tiers doc'd:
authoritative OLTP ledger (native/Formance) + ClickHouse OLAP projection over the same
o11y event stream (audit mirror — no second metering pipeline).
Tests (-race, green): Formance adapter (accrual+idempotency, debit guard+replay,
snapshot) via a fake Formance server; scope isolation (per-org never sees house);
backed-payout enforced+blocked+at-most-once; VoidPayout restores pending.
* feat(treasury): Phase 2 — Hanzo L1 (36963) ledger-root anchor (contract + luxfi/geth submit + KMS signer)
Make the off-chain books tamper-evident on the LIVE Hanzo L1 (verified running:
network/chainId 36963, hanzod-0 producing blocks, EVM at network-36963).
- contracts/TreasuryAnchor.sol: minimal immutable witness — owner-gated anchor(bytes32)
appends a timestamped root + emits Anchored; latest()/count for cheap verification.
No upgradeability, no token — one job.
- anchor_evm.go: real luxfi/geth submitter — dial → chainID/nonce/gasPrice → sign a
LegacyTx (anchor(bytes32) call when TREASURY_ANCHOR_CONTRACT set, else a 0-value
self-tx carrying the root) with types.SignTx → send → await receipt → persist. The
signer key is provisioned from KMS (KMSSecret → env TREASURY_ANCHOR_SIGNER_KEY,
ref TREASURY_ANCHOR_SIGNER_KMS_REF) — NEVER plaintext in code/manifest.
- ledger.Root/ComputeRoot: deterministic SHA-256 hash-chain over the whole journal +
reserve, shared by both backends so the anchor is backend-agnostic. A change to any
historical posting changes the root.
- POST /v1/admin/treasury/anchor submits when wired; else returns the root that WOULD
be committed + the EXACT remaining step. GET /v1/admin/treasury shows last anchored
root/tx/block + synced flag. Persisted across restart (treasury_anchor.json).
Honest status: the on-chain submit is COMPLETE + compiling + config-gated but NOT
driven live this pass — the node's external JSON-RPC is unreachable from the build
env and needs an operator to: deploy TreasuryAnchor on 36963, provision the KMS
signer (fund it), set TREASURY_ANCHOR_{RPC_URL,CONTRACT,SIGNER_KEY}. In-cluster the
cloud binary reaches hanzod-rpc-internal:9630, so it's one deploy-config away.
Builds green (cmd/cloud links luxfi/geth); tests -race green; gofmt clean.
---------
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Manage a local hanzo-engine (the `hanzoai` OpenAI + Anthropic model server)
from the canonical `hanzo` CLI:
- install: runs the canonical install.sh / install.ps1 (single source of truth
for platform detection + signature verification — no re-implementation).
- serve: launches the installed binary (`hanzoai --port P run -m MODEL`);
syscall.Exec on Unix so signals + exit code flow through.
- status: probes the local engine, reusing the /v1/models probe that
`hanzo gpu connect --serve-engine` advertises with.
Tests cover wiring, ready/unreachable status, and binary discovery.
Make the off-chain books tamper-evident on the LIVE Hanzo L1 (verified running:
network/chainId 36963, hanzod-0 producing blocks, EVM at network-36963).
- contracts/TreasuryAnchor.sol: minimal immutable witness — owner-gated anchor(bytes32)
appends a timestamped root + emits Anchored; latest()/count for cheap verification.
No upgradeability, no token — one job.
- anchor_evm.go: real luxfi/geth submitter — dial → chainID/nonce/gasPrice → sign a
LegacyTx (anchor(bytes32) call when TREASURY_ANCHOR_CONTRACT set, else a 0-value
self-tx carrying the root) with types.SignTx → send → await receipt → persist. The
signer key is provisioned from KMS (KMSSecret → env TREASURY_ANCHOR_SIGNER_KEY,
ref TREASURY_ANCHOR_SIGNER_KMS_REF) — NEVER plaintext in code/manifest.
- ledger.Root/ComputeRoot: deterministic SHA-256 hash-chain over the whole journal +
reserve, shared by both backends so the anchor is backend-agnostic. A change to any
historical posting changes the root.
- POST /v1/admin/treasury/anchor submits when wired; else returns the root that WOULD
be committed + the EXACT remaining step. GET /v1/admin/treasury shows last anchored
root/tx/block + synced flag. Persisted across restart (treasury_anchor.json).
Honest status: the on-chain submit is COMPLETE + compiling + config-gated but NOT
driven live this pass — the node's external JSON-RPC is unreachable from the build
env and needs an operator to: deploy TreasuryAnchor on 36963, provision the KMS
signer (fund it), set TREASURY_ANCHOR_{RPC_URL,CONTRACT,SIGNER_KEY}. In-cluster the
cloud binary reaches hanzod-rpc-internal:9630, so it's one deploy-config away.
Builds green (cmd/cloud links luxfi/geth); tests -race green; gofmt clean.
`hanzo gpu connect --serve-engine` advertises a local hanzo-engine (the OpenAI +
Anthropic model server on :1234) on the org fleet, alongside the existing
studio.render worker. The worker probes GET {engine-url}/v1/models, publishes the
endpoint + model list in its presence record, and prints (or with --register-provider
POSTs) the /v1/add-provider call that routes api.hanzo.ai model traffic to this GPU as
an OpenAI-compatible (Type=Local) provider.
- cli/gpu.go: --serve-engine/--engine-url/--engine-endpoint/--register-provider;
probeEngine, refreshEngine, engineAdvertisement, capabilities, provider hint;
`hanzo gpu status` shows the engine endpoint.
- clients/visor/fleet.go: byoWorker + fleetRegistration carry capabilities + engine;
GET /v1/fleet/workers advertises the endpoint (additive, omitempty).
- docs/bring-your-gpu.md: Connect (BYO) vs Deploy (cloud) -> engine.serve + studio.render.
- tests: probe/advertise/registration + full stub-cloud round-trip (no model needed).
One fleet, two job types: engine.serve (model serving) + studio.render (diffusion).
Adopt Formance as the ledger of record behind a ledger.Backend PORT, without
reimplementing double-entry: two adapters satisfy the port — the native Base/SQLite
engine (offline/default, ships the reserve fund today) and clients/treasury/formance
(a real HTTP client to the Postgres-backed Formance Ledger v2 API: world→fund accrual,
fund→payout debit, 400 INSUFFICIENT_FUND→not-backed=the overdraw guard Formance
enforces, reference→idempotency). Select by FORMANCE_LEDGER_URL — a config flip. Root
computed via a SHARED hash so the L1 anchor is backend-agnostic.
Back the growth-loop payouts: referrals/affiliates/authors now DEBIT the reserve fund
via the ONE treasury.Reserve seam before crediting the recipient wallet — fund down,
wallet up, reconciled. Not backed → honestly pending (referrals) or 402 + VoidPayout
restores pending (affiliates/authors). Idempotent by ref (at-most-once). Unmounted →
passthrough (backward-safe; existing loop suites stay green).
Scope-aware /v1/finance/* — ONE engine, three tenancy surfaces (admin/console/finance
product): tenant derived from IAM, house/reserve locked to global-admin under
/v1/admin/finance/*, per-org callers see ONLY their own org:<tenant>:* accounts.
GET /v1/finance/accounts (per-org; admin ?scope=house|?org=<t>). Storage tiers doc'd:
authoritative OLTP ledger (native/Formance) + ClickHouse OLAP projection over the same
o11y event stream (audit mirror — no second metering pipeline).
Tests (-race, green): Formance adapter (accrual+idempotency, debit guard+replay,
snapshot) via a fake Formance server; scope isolation (per-org never sees house);
backed-payout enforced+blocked+at-most-once; VoidPayout restores pending.
The platform's OWN fund/reserve accounting, one layer ABOVE the per-org commerce
credit ledger. A store-agnostic, cloud-decoupled double-entry engine
(clients/treasury/ledger) — the SEED of the native hanzoai/finance central ledger
(the Go replacement for the Formance stack) — plus a Base/SQLite adapter
(ledger/sqlstore) and the cloud client (clients/treasury).
Core (clients/treasury/ledger): accounts + balanced journal entries (Σ postings==0,
refused otherwise), ONE shared fund:reserve pool with per-program payout sinks,
revenue-share policy (bps, one place), and the reserve GUARD — a fund debit that
would overdraw is refused, atomically, so growth-loop payouts are backed capital not
unbounded minting. Zero cloud/zip/SQLite imports; persistence is the Store/Tx port,
so it lifts to hanzoai/finance as a directory move.
Surface: GET /v1/treasury (org transparency), GET /v1/admin/treasury (report +
journal + anchor), POST /v1/admin/treasury/{policy,sweep,seed,anchor} (global-admin).
treasury.Reserve(program,ref,memo,cents) is the ONE seam the 3 loops call: backed →
proceed to credit; not backed → honestly pending; unmounted → passthrough
(backward-safe). Idempotent by ref (at-most-once fund debit). ledger.Root commits the
whole journal for the Hanzo L1 anchor (Phase 2 wires the KMS-signed submit).
Tests (-race, green): double-entry balances, revenue-share accrual + per-period
idempotency, reserve guard (backed→blocked), at-most-once, concurrent no-overdraw,
admin gate, Reserve passthrough+enforced, sqlstore round-trip + tx rollback.
* feat(iam): embed IAM in the unified cloud binary as an in-process subsystem
Folds Hanzo IAM -- the identity provider serving hanzo.id (login/authorize/
token/jwks/userinfo, /v1/iam/* admin, OAuth2/OIDC, LDAP/RADIUS) -- into the
unified hanzoai/cloud binary as the LAST binary-consolidation piece
(HIP-0106: "one Go binary embeds IAM + KMS + o11y").
clients/iamsvc wraps IAM's own Beego runtime: iamserver.Init() runs the full
bootstrap without binding a listener, and web.BeeApp.Handlers is mounted
verbatim on cloud's zip.App at every prefix IAM owns (/v1/iam/*,
/.well-known/*, /login/oauth/*, /_/iam/*, /cas/*, /scim/*). No auth logic is
reimplemented -- the same controllers answer, so OAuth/OIDC semantics
(authorize clientId org-resolution, JWT audiences, SuperAdmin owner=="admin",
argon2id password hashing) are preserved byte-for-byte. Registered at order 50
(identity authority, mounts before dependents).
- go.mod: pin hanzoai/iam v1.28.12 -> v1.31.16 (latest; carries the
authorize-login org-resolution fixes #95/#96 the operator SSO chain needs).
- subsystems.go: blank-import clients/iamsvc; IAM no longer "NOT fused in".
Auth-critical middleware interactions verified: /v1/iam/* prices to 0 in
DefaultPrice (ungated -- the M2M /v1/iam/oauth/token mint is never charged);
SanitizeIdentity strips only forgeable X-User-*/X-Org-* headers, never the
Authorization bearer or iam_session_id cookie IAM's session/oauth logic reads.
Activation is STAGED via the enable-list gate: "iam" is NOT added to the live
--enable until IAM config is present in the cloud runtime and the fold is
verified (login/authorize/token/jwks + operator SSO chain). The standalone iam
pod keeps serving hanzo.id via ingress until then.
Build gate: CGO_ENABLED=0 go build ./... && go test . green. clients/iamsvc
tests prove registration (order 50) + full-path preservation through the mount.
* fix(iam): red-review — embed-mode bootstrap, fail-closed, single-replica guard
Addresses the red review of cloud#142 (mount mechanism approved; activation
blocked on standalone-only side effects in the wrapped entrypoint).
1. [HIGH] Embed-mode bootstrap. iamsvc now calls iamserver.InitEmbed (new in
iam v1.31.17) instead of the standalone Init: skips StopOldInstance
(lsof/SIGKILL — panics on distroless, kills a co-resident on shared netns),
skips LDAP/RADIUS listeners (RADIUS binds unmanaged UDP with an empty shared
secret), skips export/os.Exit, binds no listener. Standalone hanzo iam / iamd
is byte-for-byte unchanged (Init delegates to the same shared bootstrap with
every flag on). Also covers [MED] #3 — directory listeners never start
in-process.
4. [MED] Fail-closed, not fail-loud. InitEmbed returns an error (recovers
bootstrap panics); a broken/misconfigured IAM degrades THIS subsystem to a
503 fail-closed on every IAM prefix (mountFailClosed) — every co-resident
subsystem (KMS, o11y) stays up. Mirrors the KMS "no master key -> health-only"
blast-radius isolation.
2. [HIGH] Single-replica enforcement. Embedded IAM uses Beego's process-local
"memory" session store. Config.Validate now REFUSES to boot iam-enabled above
CLOUD_REPLICAS=1 (a real runtime guard, not convention); the helm chart pins
replicas=1 + injects CLOUD_REPLICAS whenever "iam" is in --enable.
5. [MED] Bump verified iam v1.31.16 -> v1.31.17. The slim-JWT change keeps every
claim cloud reads (owner, isAdmin, email, name kept; aud is a registered
claim, untouched) — IdentityMiddleware unaffected. authz v1.10.4 policy-API
swap is IAM-internal (cloud builds green, no direct use). redirect_uri
exact-match + AutoSignin CC-JWT normalization are version-skew CUTOVER gates:
version-match the standalone pod + verify registered redirect_uris are exact
before adding "iam" to the live --enable (runtime-data checklist, not code).
Tests (CGO_ENABLED=0):
- TestIAMEmbedBehindMiddlewareChain — unauth POST /v1/iam/oauth/token, /login,
jwks + /login/oauth/authorize return 2xx through the REAL SanitizeIdentity +
BillingGate chain (never 402/503); forged X-User-IsAdmin is stripped; a priced
control path is denied at zero balance (proves the gate is engaged).
- TestDefaultPriceExemptsIAM — every IAM prefix prices to 0.
- TestValidateIAMSingleReplica — iam + replicas>1 refused; 1/unset/off ok.
- TestMountFailClosed503 — the fail-soft path serves 503 on every IAM prefix.
Build+test green; standalone hanzo iam still links; helm renders replicas=1 for
iam-enabled, replicaCount otherwise. Depends on iam v1.31.17
(hanzoai/iam#feat/iam-embed-entrypoint). STILL STAGED — the standalone iam pod
serves hanzo.id until red GREEN + runtime e2e.
---------
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Addresses the red review of cloud#142 (mount mechanism approved; activation
blocked on standalone-only side effects in the wrapped entrypoint).
1. [HIGH] Embed-mode bootstrap. iamsvc now calls iamserver.InitEmbed (new in
iam v1.31.17) instead of the standalone Init: skips StopOldInstance
(lsof/SIGKILL — panics on distroless, kills a co-resident on shared netns),
skips LDAP/RADIUS listeners (RADIUS binds unmanaged UDP with an empty shared
secret), skips export/os.Exit, binds no listener. Standalone hanzo iam / iamd
is byte-for-byte unchanged (Init delegates to the same shared bootstrap with
every flag on). Also covers [MED] #3 — directory listeners never start
in-process.
4. [MED] Fail-closed, not fail-loud. InitEmbed returns an error (recovers
bootstrap panics); a broken/misconfigured IAM degrades THIS subsystem to a
503 fail-closed on every IAM prefix (mountFailClosed) — every co-resident
subsystem (KMS, o11y) stays up. Mirrors the KMS "no master key -> health-only"
blast-radius isolation.
2. [HIGH] Single-replica enforcement. Embedded IAM uses Beego's process-local
"memory" session store. Config.Validate now REFUSES to boot iam-enabled above
CLOUD_REPLICAS=1 (a real runtime guard, not convention); the helm chart pins
replicas=1 + injects CLOUD_REPLICAS whenever "iam" is in --enable.
5. [MED] Bump verified iam v1.31.16 -> v1.31.17. The slim-JWT change keeps every
claim cloud reads (owner, isAdmin, email, name kept; aud is a registered
claim, untouched) — IdentityMiddleware unaffected. authz v1.10.4 policy-API
swap is IAM-internal (cloud builds green, no direct use). redirect_uri
exact-match + AutoSignin CC-JWT normalization are version-skew CUTOVER gates:
version-match the standalone pod + verify registered redirect_uris are exact
before adding "iam" to the live --enable (runtime-data checklist, not code).
Tests (CGO_ENABLED=0):
- TestIAMEmbedBehindMiddlewareChain — unauth POST /v1/iam/oauth/token, /login,
jwks + /login/oauth/authorize return 2xx through the REAL SanitizeIdentity +
BillingGate chain (never 402/503); forged X-User-IsAdmin is stripped; a priced
control path is denied at zero balance (proves the gate is engaged).
- TestDefaultPriceExemptsIAM — every IAM prefix prices to 0.
- TestValidateIAMSingleReplica — iam + replicas>1 refused; 1/unset/off ok.
- TestMountFailClosed503 — the fail-soft path serves 503 on every IAM prefix.
Build+test green; standalone hanzo iam still links; helm renders replicas=1 for
iam-enabled, replicaCount otherwise. Depends on iam v1.31.17
(hanzoai/iam#feat/iam-embed-entrypoint). STILL STAGED — the standalone iam pod
serves hanzo.id until red GREEN + runtime e2e.
The ZAP-native trace exporter marshaled its payload via the generated
otlp/collector/trace/v1.ExportTraceServiceRequest, whose sibling
trace_service_grpc.pb.go (no build tag) drags google.golang.org/grpc into the
graph — contradicting the exporter's own contract (ZAP wire, never gRPC).
Encode the ExportTraceServiceRequest envelope directly from the grpc-free trace
messages with protowire: it is a single 'repeated ResourceSpans resource_spans
= 1', so appending each ResourceSpans under field 1 is byte-identical to the
generated marshaler (proven by the existing round-trip test, which still decodes
with the canonical collector type). go list -deps ./zaptrace now shows no grpc.
Hanzo services speak ZAP/HTTP/WS, never gRPC.
Caveat: the cloud module still pulls google.golang.org/grpc transitively via
hanzoai/ai (sibling-owned), hanzoai/o11y (embedded SigNoz — intrinsically an
OTLP/gRPC collector) and hanzoai/base (GCS gRPC transport). Not removable by a
cloud-local change; tracked separately. go.sum: incidental tidy prune of stale
vfs/age checksums.
Folds Hanzo IAM -- the identity provider serving hanzo.id (login/authorize/
token/jwks/userinfo, /v1/iam/* admin, OAuth2/OIDC, LDAP/RADIUS) -- into the
unified hanzoai/cloud binary as the LAST binary-consolidation piece
(HIP-0106: "one Go binary embeds IAM + KMS + o11y").
clients/iamsvc wraps IAM's own Beego runtime: iamserver.Init() runs the full
bootstrap without binding a listener, and web.BeeApp.Handlers is mounted
verbatim on cloud's zip.App at every prefix IAM owns (/v1/iam/*,
/.well-known/*, /login/oauth/*, /_/iam/*, /cas/*, /scim/*). No auth logic is
reimplemented -- the same controllers answer, so OAuth/OIDC semantics
(authorize clientId org-resolution, JWT audiences, SuperAdmin owner=="admin",
argon2id password hashing) are preserved byte-for-byte. Registered at order 50
(identity authority, mounts before dependents).
- go.mod: pin hanzoai/iam v1.28.12 -> v1.31.16 (latest; carries the
authorize-login org-resolution fixes #95/#96 the operator SSO chain needs).
- subsystems.go: blank-import clients/iamsvc; IAM no longer "NOT fused in".
Auth-critical middleware interactions verified: /v1/iam/* prices to 0 in
DefaultPrice (ungated -- the M2M /v1/iam/oauth/token mint is never charged);
SanitizeIdentity strips only forgeable X-User-*/X-Org-* headers, never the
Authorization bearer or iam_session_id cookie IAM's session/oauth logic reads.
Activation is STAGED via the enable-list gate: "iam" is NOT added to the live
--enable until IAM config is present in the cloud runtime and the fold is
verified (login/authorize/token/jwks + operator SSO chain). The standalone iam
pod keeps serving hanzo.id via ingress until then.
Build gate: CGO_ENABLED=0 go build ./... && go test . green. clients/iamsvc
tests prove registration (order 50) + full-path preservation through the mount.
The THIRD growth loop next to referrals (one-time credit) and affiliates
(partner commission): pays open-source AUTHORS a royalty on the metered platform
spend of orgs who DEPLOY their projects on Hanzo. Mirrors clients/affiliates
exactly — one SQLite store, server-side tenant isolation, one Mount (HIP-0106),
the SAME commerce ledger path (a credits payout is a grant, tag grant:author),
and an at-most-once accrual latch.
Flow: connect GitHub (IAM-linked account or supplied login) → verify repo
ownership (OAuth admin-check OR a hanzo.json verify-code file) → a deploy of a
verified author repo by ANY org is recorded (provenance) → sweep accrues 5% of
that org's month-to-date spend, at-most-once per (author, deploying-org, period),
self-deploys excluded → staff pay out as credits (real grant) or cash (record-only),
never exceeding pending.
Surface: GET /v1/authors, POST /v1/authors/{connect,repos/verify,deploys/record};
GET /v1/admin/authors, POST /v1/admin/authors/{sweep,:id/approve,:id/suspend,:id/payout}.
10 tests, all -race green: repo canonicalization, both verify methods, deploy
attribution + idempotency, spend×share accrual + at-most-once, lazy dashboard
sweep, credits-one-grant/cash-record-only/pending-guard payout, admin gate.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Mirrors clients/referrals: one SQLite store, server-side tenant isolation, one
HIP-0106 Mount, admin surface global-admin-gated + enveloped for the console
proxy. Affiliates earn an ONGOING commission (default 20%) on the metered spend
of the customers they refer — the recurring, partner-revenue growth loop beside
referrals' one-time both-sides credit.
- apply (org) -> status applied; staff approve mints the code (vanity opt-in,
uniqueness-enforced, else a derived slug) + sets the rate.
- attribute (?aff capture) records referred_org->affiliate (first-touch, one per
referred org, self blocked; approved affiliates only).
- accrual sweep: commission = referred org spend this period x rate, latched
at-most-once per (affiliate, referred_org, period) in one txn; also lazy on the
affiliate's own dashboard read.
- payout: a credits method issues a commerce grant (tag grant:affiliate); cash
methods are record-only; can never exceed pending (accrued - paid), reserved
atomically before any grant.
Tests (go test -race, 9 green): apply->approve, vanity uniqueness (409),
accrual = spend x rate, idempotent-per-period sweep, payout-as-credits issues one
grant + cash record-only + pending guard, admin gate 403, attribution
self/unknown/first-touch, Mount.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
* feat(referrals): native /v1/referrals viral loop over the commerce ledger
Per-org referral program mirroring clients/crm's structure (one SQLite store,
server-side tenant isolation, HIP-0106 Mount). Grants promo credit through the
SAME commerce deposit path as clients/admin.grantCredit (trial/Credit bucket,
tag grant:referral).
- Stable deterministic referral code per org (base32 of a hash of the org id) +
a white-labeled ?ref link; persisted directory for O(1) reverse lookup.
- POST /v1/referrals/claim: record referrer<->referee (referee = validated
caller), status signed_up. Self-referral blocked, one-per-referee idempotent
(first-touch wins).
- Qualify signal = referee metered spend (honest 'actually used the product').
On qualify, grant BOTH sides: referrer +$10, referee +$5. At-most-once via a
credited_at latch — no sweep and no concurrent read can double-pay.
- Trigger: lazy on the referrer's GET /v1/referrals + POST /v1/admin/referrals/
sweep (cron path). GET /v1/admin/referrals directory, both global-admin gated.
- Constants (bonus amounts + ledger tag) in one place. Commerce behind an
interface for testable double-grant/idempotency proofs.
Tests: code derivation, self-ref block, idempotent claim, qualify->double-grant
with balances moving through the (fake) ledger, at-most-once idempotency, lazy
qualify on read, admin gate + directory, real Mount. All green (go test -race).
* feat(referrals): envelope the /v1/admin/referrals surface for the console admin proxy
---------
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
cloud's internal/org was the ORIGIN of the HA-SQLite machinery; it's now promoted to the
shared hanzoai/vfs/replica lib that every service adopts. Delete the duplicated impls
(replica.go Replicator+Store+DB+DBPath, owner.go Member+Owner+IsOwner+Replicas+HRW) and
re-export them as aliases (shared.go). cloud-specific pieces stay: membership.go (live IAM
source), cipher.go (KMS envelope — already satisfies replica.Cipher), vfsstore.go (Store over
deps.VFS, now using the exported replica.Version). One and one way: ONE Replicator + election,
in vfs/replica, used by cloud AND visor. Builds + org tests green (vfs v0.6.2).
Zen models were capped at the 4096 fallback in getContextLength, so every
console chat (grounded assistant ~4190-token system prompt) 402'd
'exceeds maximum token count: 4096'. v1.800.9 special-cases the zen* prefix
to 131072. Fixes the P0 console-chat gate.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The embedded o11y runtime's OTel instrumentation defaults to a Prometheus pull
reader bound to 0.0.0.0:9090 (pkg/instrumentation) — the SAME port as cloud's
health listener (CLOUD_HEALTH_LISTEN=:9090). Activating the embed therefore made
the whole cloud process crash-loop with 'listen tcp :9090: bind: address already
in use' (verified on the canary), taking down all of api.hanzo.ai — a listener the
standalone o11y pod never contended for.
buildEmbeddedHandler now defaults O11Y_INSTRUMENTATION_METRICS_ENABLED=false (via
setenvDefault, operator-overridable to a free port) before construction. Cloud owns
process-level observability (exports its own OTel telemetry), so the embed serves
/v1/o11y in-process without a second metrics listener. Extracted the env defaults
into applyEmbedEnvDefaults + TDD (guard + operator-override).
Verified on the cloud-unified-canary: with this default the .104 embed goes Ready
and serves /v1/o11y in-process (health 200); without it the pod crash-loops on :9090.
CGO=0 go build/test green.
Refactor clients/o11y/embed.go onto o11y v1.5.0's shared builder community.NewServer
+ community.NewConfig — the EXACT construction the standalone o11y pod runs — so
the in-process runtime cannot drift from the pod's auth (pkg/identn/iamidentn,
Hanzo IAM gateway-header identity). Collapses the duplicated ~70-line signoz.New
factory list (drift risk) to one call. Enable signal now reads the flat operator
knob O11Y_DATASTORE_DSN (what the pod sets), falling back to the structured
O11Y_TELEMETRYSTORE_DATASTORE_DSN.
Exempt liveness/readiness paths from gate(): the runtime serves them without
identity (k8s probes pass that way), so gating them only breaks unauthenticated
health probes (admin System Health CLOUD_O11Y_HEALTH_URL, the external o11y.*
hosts) without protecting anything. Data routes stay gated (RED forge test still
403s /v1/o11y/api/v1/query_range).
go.mod: o11y v1.4.1 -> v1.5.0 (identical go.mod hash — no new transitive deps).
Build-gate: CGO_ENABLED=0 go build ./... = 0, go test . = ok, go test ./clients/o11y = ok.
The #133 embed pinned o11y v1.3.13, whose runtime authenticates via o11y-native
JWT (tokenizer.GetIdentity on the Authorization bearer). The live gateway-header
traffic the standalone o11y:0.2.0 pod serves — identity injected as X-Org-Id/
X-User-Id/X-User-Email by the gateway — would 401 against that. So activating the
v1.3.13 embed could not replace the pod.
This repoints the embed to the MAIN o11y line (v1.4.1), which resolves identity
through the IdentN resolver's iamidentn provider (default-enabled) from those
gateway session headers, with iamauthz (Hanzo IAM Casbin) for authorization —
the SAME auth model as the running pod. Gateway-header traffic authenticates
(200), not 401.
- clients/o11y/embed.go: build the runtime via pkg/signoz.New with the SAME
provider factories the standalone cmd/community server uses (noop zeus,
licensing, gateway, auditor, meterreporter; iamauthz; ClickHouse
telemetrystore; sqlite sqlstore; IdentN to iamidentn), then app.NewServer to
server.PublicHandler (new accessor, o11y v1.4.1). runtime.Start runs the
registry background services (incl. the ruler/alert-rule-manager)
non-blocking; we never call server.Start (cloud owns its HTTP listeners; OpAMP
stays out-of-process). Gate/proxy-fallback structure (clients/o11y/o11y.go) is
unchanged: still O11Y_TELEMETRYSTORE_DATASTORE_DSN-gated, still fail-soft to
the reverse proxy.
- go.mod: hanzoai/o11y v1.3.13 to v1.4.1 (main line, iamidentn). Drop the stale
replace prometheus/alertmanager to hanzoai/alertmanager v0.28.2 — it forced
o11y's code onto the old fork whose api/v2 returns hanzoai/common types that
clash with o11y v1.4.1's upstream prometheus/common structs. o11y v1.4.1 (and
the pod) build against upstream prometheus/alertmanager v0.31.1; cloud has no
direct alertmanager import, so it now matches.
Telemetry backend (ClickHouse datastore StatefulSet, cluster insights) is
untouched — the embedded runtime queries it over ClickHouse-native :9000.
Build-gate: CGO_ENABLED=0 go build ./... OK; go test ./clients/o11y/... OK; vet OK.
v1.49.0 adds the durable workflow primitives social-orchestrator needs to run
on cloud's embedded gated engine (ServeGated :9999): signal-to-running-workflow
re-dispatch, continueAsNew, startChild, typed search attributes, workflowId
conflict policy, and the signalWithStart wire fix. No cloud code change — the
embedded engine + gated listener pick up the fixes on rebuild.
Constructs the ONE hanzoai/o11y runtime IN-PROCESS (clients/o11y/embed.go) —
the SAME bootstrap the standalone cmd/server runs (o11y.New with its provider
factories -> app.NewServer -> server.PublicHandler) — and installs it via
o11y.SetHandler, so /v1/o11y/* is served by THIS binary against the ClickHouse
`datastore` (StatefulSet, cluster insights) instead of reverse-proxying a
standalone o11y Deployment. The standalone o11y pod can now retire; the
ClickHouse datastore stays as the telemetry backend.
- clients/o11y/embed.go: buildEmbeddedHandler wires telemetrystore (ClickHouse/
datastore), sqlstore (sqlite under cloud's data root), querier, dashboards,
alerts; starts the registry services + the alert rule manager (StartBackground).
Enabled by O11Y_TELEMETRYSTORE_DATASTORE_DSN (the DSN is the one knob).
- o11y.go Register callback: prefer the in-process runtime; fall back to the
reverse proxy when the embed is disabled (no DSN) or fails to init — fail-soft,
zero downtime. Proxy handler + gate + tests retained for the fallback path.
- Bump hanzoai/o11y v1.3.12 -> v1.3.13 (adds Server.PublicHandler + StartBackground).
- Drop the stale `replace gorilla/mux => containous/mux`: it was a copied Traefik
replace block; Traefik is not in the graph and nothing calls the containous API,
but the fork lacks mux.MiddlewareFunc that o11y's otelmux needs. Standard
gorilla/mux v1.8.1 satisfies every consumer.
Deferred (reported, not faked): OpAMP collector management (a second websocket
listener) is not started in-process — telemetry ingest continues on the existing
collector->datastore path. Build/test gate is CGO_ENABLED=0 (as prod ships): o11y
+ hanzoai/sqlite resolve to a single modernc sqlite driver registration.
Consolidation: kill the standalone tasksd pod by running its consumers on cloud's
in-process embedded engine. After Embed wires the loopback (ungated, in-process
ai-ingest) listener, call emb.ServeGated(ctx, 9999, validator) to expose the SAME
engine cluster-wide under mandatory identity gating.
RequireIdentity: every request on :9999 must carry an IAM auth_token, validated
against {IAMIssuer}/v1/iam/.well-known/jwks (HIP-0111) and org-scoped to its owner --
the same trust anchor as the HTTP SanitizeIdentity boundary. The loopback dialer for
ai-ingest is untouched (127.0.0.1:19999, ungated, cloud's own trust boundary).
Fail-soft: a missing IAMIssuer or a bind failure logs and leaves the gated surface
down without disturbing ai-ingest. 9999 mirrors the port the retired tasksd exposed,
so a consumer repoint changes only the host (tasks.hanzo.svc -> cloud.hanzo.svc).
Depends on hanzoai/tasks#8 (ServeGated + identity over ZAP). Pinned here to that
branch's commit; repin to the tagged release once #8 merges. universe adds the :9999
Service port.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Consolidation: kill the standalone tasksd pod by running its consumers on cloud's
in-process embedded engine. After Embed wires the loopback (ungated, in-process
ai-ingest) listener, call emb.ServeGated(ctx, 9999, validator) to expose the SAME
engine cluster-wide under mandatory identity gating.
RequireIdentity: every request on :9999 must carry an IAM auth_token, validated
against {IAMIssuer}/v1/iam/.well-known/jwks (HIP-0111) and org-scoped to its owner --
the same trust anchor as the HTTP SanitizeIdentity boundary. The loopback dialer for
ai-ingest is untouched (127.0.0.1:19999, ungated, cloud's own trust boundary).
Fail-soft: a missing IAMIssuer or a bind failure logs and leaves the gated surface
down without disturbing ai-ingest. 9999 mirrors the port the retired tasksd exposed,
so a consumer repoint changes only the host (tasks.hanzo.svc -> cloud.hanzo.svc).
Depends on hanzoai/tasks#8 (ServeGated + identity over ZAP). Pinned here to that
branch's commit; repin to the tagged release once #8 merges. universe adds the :9999
Service port.
HIGH-2: principal.ValidatedProject(c) (project, validated) — returns false
today (X-Project-Id is a caller-chosen label, not claim-bound), so the edge
gate + resource meter pass ProjectValidated=false and commerce degrades
project-scoped hard caps to soft. ONE lever to harden when IAM mints a
project claim. MED-4: DenyResource renders ErrSpendCapExceeded -> 402
spend_cap_exceeded (was 503). INFO-7: canonicalService unifies the edge
service label with the resource provider (ml/visor->compute, agents->agent,
security->security.scan) so a cap binds on both surfaces. Bump metering
v0.1.3 -> v0.1.4. Tests green (incl DenyResource spend_cap).
Adds TEAM_PUBLIC_URL / PUBLIC_ORIGIN config: callbackOrigin() returns the
configured public origin (e.g. https://hanzo.team) for the OAuth redirect_uri
instead of the request Host, so cloud emits the registered public callback even
behind the gateway (where the request Host is the internal cluster service).
Unset = unchanged (falls back to originOf). Lets hanzo.team route through the
gateway UNIFORMLY like api.hanzo.ai — removes the need for the temporary
direct-to-cloud edge route.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
BillingGate uses metering.AuthorizeVerdict (funds+cap, one round trip):
renders a distinct 402 spend_cap_exceeded (scope/cap/spent) and sets
X-Spend-Warn at the soft threshold; gates on the request price. New ONE
ScopeRateLimit middleware composes zip/middleware.RateLimit per-scope
(org/project/service), dynamic rpm from commerce (short-TTL cache,
fail-open), 429 + X-RateLimit-* — wired after identity, before billing.
ResourceMeter threads project + service(=provider) so resource creation
is scope-gated too. Scope always from the validated principal. Bump
metering v0.1.2 -> v0.1.3. Money-path tests green (402/warn/429/isolation).
pickVFSClient returns an S3-backed types.VFSClient (clients/s3vfs.go) when
S3_ADMIN_ACCESS_KEY/SECRET_KEY are set, else DisabledVFS (R-7 fail-closed
preserved). Reuses the SAME s3admin.Admin construction as clients/s3 (DRY, one
credential path). Put/Get/Delete over the shared 'team-blobs' bucket, per-tenant
key prefix (files.go builds team/blobs/<verified-org>/<ws>/<blobId>). S3 NoSuchKey
maps to types.ErrBlobNotFound (honest 404/idempotent-204); any other S3 error →
502 fail-closed (never a dishonest 404). Bucket create-if-absent self-heals a boot
blip. Red-reviewed SHIP. This is the repoint gate: hanzo.team avatars/attachments
now work off cloud.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Ports hanzoai/team-go into the unified cloud binary as a zip-native subsystem
(order 138, /v1/team/*): account (IAM OAuth bridge + workspaces/members on SQLite),
transactor (Huly wire over wsx, serverVersion 0.6.0 preserved), bots-as-members
(in-process agents.ListForOrg → Employees, removal-reconcile), files (FrontStorage
contract, org+workspace-membership scoped, byte-derived content-type allow-list).
Security (Red-reviewed, all closed): fail-closed SERVER_SECRET degrade-health-only
(never crashes the binary/CI smoke-boot), token exp/nbf, seg() traversal guard,
setCookie verify, cross-tenant blob isolation, VFSClient.Delete fail-closed (deps.VFS
never nil, R-7). Supersedes the stale clients/team a parallel branch swept onto main.
Real deps.VFS wiring (avatars) follows in the next patch (.97) before the hanzo.team
front repoint, so nothing regresses. Migration + repoint + rip of the standalone
team-go Deployment are the remaining cutover steps.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
- POST /v1/admin/customers/:org/credit gains `source` (trial|prepaid). A staff
comp defaults to TRIAL (non-cash Credit bucket, grant:admin tag → billing/bucket
DepositKind Credit); only explicit "prepaid" mints real money (admin-grant tag →
Prepaid). Fail-closed: unknown→trial, so a comp never silently becomes payout-able
cash. source recorded in the audit before/after + response.
- grantCredit refactored to a shared applyGrant core (ONE credit-write path).
- NEW GET /v1/admin/grants — the credit-grant ledger across all orgs, projected
from the tamper-evident audit trail (action admin.customer.credit): org, amount,
source, reason, staff actor, date, txid, result. Honest-empty without a local
audit store.
- NEW POST /v1/admin/grants — issue a grant to any org from the operator Grants
view (org in body), funneled through the SAME applyGrant core.
- Both global-admin gated (s.guard). grantTag unit-tested.
go build/vet/test ./clients/admin green.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
- POST /v1/admin/customers/:org/credit gains `source` (trial|prepaid). A staff
comp defaults to TRIAL (non-cash Credit bucket, grant:admin tag → billing/bucket
DepositKind Credit); only explicit "prepaid" mints real money (admin-grant tag →
Prepaid). Fail-closed: unknown→trial, so a comp never silently becomes payout-able
cash. source recorded in the audit before/after + response.
- grantCredit refactored to a shared applyGrant core (ONE credit-write path).
- NEW GET /v1/admin/grants — the credit-grant ledger across all orgs, projected
from the tamper-evident audit trail (action admin.customer.credit): org, amount,
source, reason, staff actor, date, txid, result. Honest-empty without a local
audit store.
- NEW POST /v1/admin/grants — issue a grant to any org from the operator Grants
view (org in body), funneled through the SAME applyGrant core.
- Both global-admin gated (s.guard). grantTag unit-tested.
go build/vet/test ./clients/admin green.
Align cloud to the target: KMS is the embedded luxfi/kms (clients/kms) alongside
embedded IAM; the last external hanzoai/kms dependency is removed.
- go.mod: bump github.com/luxfi/kms v1.11.6 → v1.11.8 (match the deployed image);
remove github.com/hanzoai/kms/sdk/go v1.1.1; go mod tidy.
- clients/mpcseal (NEW): the minimal client-side-CEK sealing client for the
SEPARATE luxfi/mpc node ring — a faithful, behavior-identical inline of the
subset of the former hanzoai/kms/sdk/go that clients/fleet + clients/provisioning
use (NewClient/Unlock/Set/Get/Delete + Argon2id→HKDF→AES-256-GCM). luxfi/kms has
no drop-in equivalent (its pkg/ is server/ZAP/store, not a Vault client), so
inlining the used subset is the minimal correct change that removes the external
dep without altering the wire protocol or trust model. Drops the HPKE Wrap/Unwrap
the callers never used.
- clients/{fleet,provisioning}: swap import path only; call sites untouched.
- clients/kmssvc/login.go + docs/consolidation.md: comment/table refs → luxfi.
Verified: go build ./... = 0, go vet = 0, gofmt clean, clients/provisioning tests
pass. go.mod + go.sum carry zero hanzoai/kms references.
Follow-up (separate, tested change): fold fleet/provisioning sealing into cloud's
embedded deps.KMS once types.KMSClient gains Delete + verified against the live MPC
ring — one KMS surface. Once the universe KMS-collapse PR merges, the deprecated
hanzoai/kms repo/package can be archived.
The gateway now mints X-Project-Id (an org SUB-SCOPE) alongside X-Org-Id.
Thread it through the keyed surfaces, backward-compatibly — the default
project ("default", or an absent header) resolves to today's exact keys,
so existing single-project tenants are byte-identical.
- principal.Project(c): the ONE read accessor, mirroring c.Org() (zero-copy
header read, cloned on retain). Defaults to DefaultProject when the header
is empty. principal.DefaultProject / IsDefaultProject own the default-scope
semantics in one place (shared contract value with iamauth.DefaultProject).
- fleet: registry refs shard by project via the ONE scopeRef seam —
"<org>/fleet/clusters" for the default project, "<org>/<project>/fleet/
clusters" for a non-default one (index, sealed kubeconfig, cache key).
- ml: tenant namespace is "ml-<org>" for the default project and
"ml-<org>-<project>" for a non-default one; both org and project are
validated against strict DNS-label regexes (no lossy fold) and the composed
label is length-checked against the 63-char ceiling, keeping the
(org, project) -> namespace map injective. A hanzo.ai/project attribution
label is stamped for non-default projects.
- visor BYO fleet + ml federation resolve project via principal.Project.
Billing stays keyed on the paying org (a project has no separate prepaid
balance); project is isolation + attribution, not a billing key.
luxfi/age v1.5.0 was upstream-retagged (transient files GC'd from the tag
tree), so the tag's zip content on the origin no longer matches the h1 hash
recorded in go.sum. Cold-cache builds (fresh CI, empty GOMODCACHE) fail with:
verifying github.com/luxfi/age@v1.5.0: checksum mismatch
SECURITY ERROR
v1.5.1 dereferences the same commit, is immutable, and is sum.golang.org
verified (h1:Gj8iHMMi0lGkKT/mlXV2HVBr2m3vt2v0eKVsTMTtAQM=). Surgical: age
require + go.sum only. go mod verify clean.
* feat(crm): startup-program applications resource (intake + AI screen + pipeline)
Public unauthenticated intake POST /v1/crm/applications (rate-limited + honeypot)
writes a dedicated crm_applications record (all fields in metadata JSON), a
best-effort CRM Company+Contact projection, and kicks off an AI screen via the
gateway (score / tier1 / suggested credits / summary / draft reply) that
auto-advances applied->screened. Staff GET/PATCH drive a stage machine
(applied->screened->qualified->credits-offered->onboarded, +rejected w/ reason).
Non-fatal if the LLM is unavailable.
* test(crm): startup applications — intake, honeypot, idempotency, AI screen, stage machine
10 tests: public intake creates application+CRM projection with all fields in
metadata; honeypot drop; validation; idempotent resubmit; end-to-end AI screen
with a fake gateway (score/tier1/credits/reply + auto-advance applied->screened);
non-fatal screen on gateway error; staff PATCH stage machine (advance/skip-block/
reject-requires-reason); pure canTransition + parseScreen + detectTier1.
---------
Co-authored-by: hanzo-dev <dev@hanzo.ai>
PR #124 embedded the last-known-good admin-tasks build as a de-risking
fallback. This replaces it with a FRESH build of current admin-tasks HEAD,
built from source in a combined gui+admin workspace (hanzogui@7.3.x +
@hanzogui/admin@7.3.0 workspace-linked; @hanzogui@7.3.x is unpublished, so it
must build inside that workspace — see clients/tasksvc/ui/README.md).
Same contract (base=/_/tasks/, api=/v1/tasks); full component parity
(namespaces/workflows/schedules/batches/deployments/activities/nexus/history).
Embed tests still assert the real bundle (not the placeholder).
One and one way: BYO k8s / BYO-GPU / bare-metal attach now lives on the SAME fleet
surface as managed clusters (visor /v1/clusters), not a parallel /v1/ml/clusters.
- clients/fleet: the shared per-org BYO-cluster Registry — kubeconfig sealed in the
org's KMS, validated by reaching the cluster (node + nvidia/amd GPU inventory),
tenant-scoped by the ZAP-propagated X-Org-Id. ONE source of truth.
- visor: POST /v1/clusters (attach) + DELETE /v1/clusters/:id (detach), and BYO
clusters MERGE into GET /v1/clusters beside managed ones. Nominal management fee
(rides the compute-fee config — no bespoke env var; customer brings the compute).
- ml: deleted the parallel /v1/ml/clusters; dynForOrg federates ML serving onto the
org's registered cluster via the shared registry (home client when none).
Builds + vet + tests green.
clients/tasksvc served /_/tasks from github.com/hanzoai/tasks/ui, whose
ui/dist is an empty 'No UI build present' placeholder — so tasks.hanzo.ai
still routed to the standalone tasks-ui pod (a Temporal-Web-UI fork).
cloud is the ONE process that serves tasks.hanzo.ai (durable.go's embedded
engine + /v1/tasks surface), so cloud now owns the UI embed too: a local
clients/tasksvc/ui package bakes the real admin-tasks SPA build (base=/_/tasks/,
api=/v1/tasks) into the binary via //go:embed. One binary, one origin, the
real UI — which lets the tasks-ui Deployment/Service/CR be retired.
Tests prove the embedded bundle is the real SPA (not the placeholder), the
SPA deep-link fallback, immutable asset caching, and GET-only.
Companion security test to the observe subsystem: proves a /v1/o11y/logs
request lands on the org-scoped handler (order 44), never falling through to
the unscoped hanzoai/o11y reverse-proxy wildcard (order 70) that would bypass
tenant scoping (attack #4).
Validated the handler SQL against the live signoz_traces schema: response_status_code
is LowCardinality(String), so a raw >= 500 raises NO_COMMON_TYPE and asInt64 on it
yields 0. Wrap with toInt32OrZero() in the RED errs count and the request-log status,
matching the verified live query (real per-org buckets returned).
The fe3a5fd agents-metering refactor split MeterUsage out of Meter and
dropped the Model:kind write, so EVERY per-product debit (functions/invoke,
s3/op, provisioning, ml, tracker, automations, security) recorded an empty
model — losing per-item revenue attribution in the commerce ledger. Restore
the one-place mapping in Meter (all 8 resource callers flow through it).
Also fix two stale test doubles that read the retired X-IAM-Org-Id header;
commerce reads X-Org-Id only (same $0-revenue class as the admin.go fix), so
they saw an empty org. Full suite: 58 ok, 0 fail.
Register a machine's GPU into the cloud fleet with one command and see it on
the console's existing Machines + GPUs pages, tagged provider=byo.
- clients/visor: a BYO worker is a heartbeating presence activity in the org's
`fleet` tasks namespace (cloud.EmbeddedTasks). fleet.go reads it and folds it
into the SAME machineView/gpuView the console renders (provider=byo,
location=on-prem, gpu model + VRAM, online/offline by heartbeat), plus a raw
GET /v1/fleet/workers. /v1/machines and /v1/gpus union Visor's inventory with
the BYO workers and degrade gracefully (BYO stays visible if Visor is down).
- cli: `hanzo gpu connect|status|disconnect` — reuses the `hanzo login`
IAM token (org from its claims; token auto-refresh), detects GPUs via
nvidia-smi, registers + heartbeats the fleet presence record, and runs an
outbound worker loop claiming from `gpu-jobs` (pluggable handlers: echo,
studio.render→local ComfyUI). --daemon installs a systemd --user unit.
- go.mod: hanzoai/tasks v1.46.0 → v1.47.0 (the claim + lease-reaper surface).
E2E (z@hanzo.ai): connect → GB10 'spark' shows provider=byo on
/v1/fleet/workers + /v1/machines + /v1/gpus → echo job claimed + completed.
launchBot bound an agent *name* but never created that agent, so
messageBot's in-process run (/v1/agents/:agent/run -> Resolve) 404'd
"agent not found" — a launched bot could not be messaged.
launchBot now create-if-absent's the bound agent via the SAME
POST /v1/agents the console uses (one create path, forwarding the
caller's validated identity -> org-scoped, IDOR-safe), BEFORE launching
the metered machine so a bad request (e.g. a non-catalog model) 400s
before anything is provisioned. Idempotent: an existing agent (409) is
reused. An omitted model takes the deployment default
(deps.AIDefaultModel, a valid catalog model) threaded from config — no
hardcoded model id.
Also add create/update-time model validation: a client-supplied model
outside the gateway's served catalog is a clean 400 (via the optional
types.ModelLister the real gateway client implements) instead of a
confusing run-time 502; fail-open when the catalog can't be enumerated.
Tests: agents model-validation + default (real store); httpAI.Models
against a fake gateway; visor launch->auto-create->message->resolve E2E
incl. the before/after 404->200 gap proof, idempotency, and bad-model
fail-fast (no machine provisioned).
notify's send surface now reads provider credentials EXCLUSIVELY from cloud's
embedded KMS (cloud.Deps.KMS) at the org-scoped, rotatable ref
orgs/<org>/notify/<svc>/<key> — the same /orgs/<org> namespace
clients/integrations uses, so a cred is seedable + rotatable via
POST /v1/kms/orgs/:org/secrets with no operator-injected env Secret and no
restart. The org is the VALIDATED principal's tenant, never a client header.
Removes the env-first fallback (envCreds/envFirst + the os import): no secret
is ever read from the environment, hard-coded, or logged. A missing key leaves
the value empty and constructProvider fails closed.
Tests rewritten to inject a fake KMS (no env), plus a per-org isolation test
and a regression that creds() ignores the legacy TWILIO_* env entirely.
Pulls hanzoai/ai#71: the RAG default embedder (object/init.go seed) now points at
the Hanzo gateway (CLOUD_AI_BASE_URL / CLOUD_AI_API_KEY, model text-embedding-qwen3)
instead of an empty ProviderUrl that hit api.openai.com directly. A server-side
embed to api.openai.com from in-cluster crawled ~180-210s and then failed, so RAG
ingest (/v1/rag/embed) hung AND the Qdrant vector collection was never created
(writeDocsToVector sample embed timed out before ensureVectorCollection ran).
Gateway embeddings are <1s (proven live). The seed self-heals an existing
api.openai.com-direct default to the gateway on boot, so this deploy converges the
live default-embed provider with no manual console repoint.
The console (console2 WebSearch module) reaches search through the /cloud proxy
with a signed-in USER BEARER, not the shared X-API-Key. searchGuard required the
key on every call (F2 hardening), so the console got 503/401 — "backend not
initialized" — even though searxng+crawl are deployed and the upstream defaults
are correct.
Reconcile to the ONE-WAY gate the rest of the /v1 data plane uses: at the zip
layer, a request with a validated principal (principal.Validated — X-User-Id
minted by the identity middleware from a verified JWT) proxies straight to
SearXNG; a request with NO principal falls to the unchanged key-based searchGuard
(the hanzo.chat server path). A caller with neither is still refused, so F2 (no
open metasearch proxy) holds — proven by TestSearchNoPrincipalNoKeyRefused.
searchGuard (net/http) is untouched; its 503/401 tests stay green. New coverage:
TestSearchValidatedPrincipalBypassesKey (console bearer, key unset -> 200),
TestSearchNoPrincipalNoKeyRefused (anonymous, key unset -> 503).
agent-runner mints M2M inference token from in-cluster IAM (fixes /v1/agents/:ref/run 502). Forward-integrated with main (#118 admin-guard audience). Reconciles live sha-b9639df onto a semver release.
cloud-api SanitizeIdentity validates the forwarded IAM bearer against
defaultJWTAudiences before granting global-admin (owner==adminOrg). The
admin.hanzo.ai guard is client hanzo-admin-guard, so its tokens carry
aud=hanzo-admin-guard, which was missing from the allowlist -> the bearer
failed validation, resolved anonymous, and the SuperAdmin gate read false
-> 403, even though the token owner IS admin.
Append the guard client_id (forwards-only, mirrors gateway iamauth). Admin
authority still requires owner==adminOrg, so no widening. Pairs with
hanzoai/gateway audience fix.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
pickAIClient built the M2M token URL from cfg.IAMIssuer (=https://hanzo.id), a
Cloudflare-fronted host. In-cluster the runner's server-side POST to
https://hanzo.id/v1/iam/oauth/token 403s with CF edge error 1006, so the
oauth2 client-credentials fetch fails and EVERY POST /v1/agents/:ref/run 502s:
'cloud: chat completion: oauth2: cannot fetch token: 403 Forbidden / 1006'.
New aiM2MTokenURL resolves the token endpoint split-horizon, mirroring the KMS
login-broker (clients/kmssvc) exactly — one policy, no drift:
1. CLOUD_AI_IAM_TOKEN_URL override
2. in-cluster IAM_URL (already wired to http://iam.hanzo.svc for JWKS)
3. public IAMIssuer fallback (single-process deploys)
IAMIssuer stays https://hanzo.id for JWT iss-validation (untouched). The chat
base URL is pointed in-cluster via CR env CLOUD_AI_BASE_URL (universe).
Proven in-cluster: mint token from iam.hanzo.svc + chat to gateway.hanzo.svc
both 200 (real completion). Unit test pins the 3-branch resolution order.
Pulls in the ai fix that closes the hz_ widget-key free-inference hole: widget
keys now bill the OWNER ORG (object.WidgetKeyOwner), so reserveBudget +
recordUsage + the balance gate all engage instead of running free/unmetered.
Bounded to the restricted widget model set + token cap; fail-secure when a widget
key is unattributable.
The per-tenant KMS secret-sync login broker derived its IAM token-exchange URL from the public issuer (hanzo.id), which Cloudflare 403s for in-cluster server-side POSTs → the sync could never authenticate. Prefer in-cluster IAM_URL (+ CLOUD_KMS_IAM_TOKEN_URL override), fall back to issuer. Unblocks PaaS per-tenant secret env (proven: git-built ai-demo app deployed on maxpower).
Mounts /v1/notify/{send,send/sms,send/email,health} natively in-process as the
cloud subsystem "notify" (order 139) — the native, in-process replacement for
the standalone notifyd (github.com/hanzoai/notify) Deployment.
notifyd's ONLY production consumer is Hanzo IAM's OTP send
(POST /v1/notify/send?sync=true, event=iam.otp_sent), and the live tenant's
template/provider/event tables are empty, so this folds exactly that contract
and nothing more. It reuses notifyd's OWN public provider packages
(service/{twilio,twilioemail,plivo,mail}) and wire types (pkg/types) — no
duplication of provider plumbing; only the internal-only cred->constructor glue
is mirrored.
Security: unlike the ClusterIP-internal notifyd (which trusted a raw X-Org-Id),
/v1/notify/send is reachable via the public gateway here, so it gates on a
VALIDATED principal and derives the org from principal.Tenant — the same
trust-boundary move clients/auto makes. Credentials come from env (the
KMS-synced notify-twilio Secret) and KMS via cloud.Deps.KMS; none is hard-coded
or logged. Ships a built-in iam.otp_sent template so the fold is strictly more
available than notifyd is today (whose empty store would 400 an OTP send).
Sync-only: the Temporal notify-send async plane is intentionally NOT folded;
async (no ?sync=true) returns 503, exactly as notifyd does without a worker.
Build-gated: go build ./... green; go test ./clients/notify/... green; gofmt/vet
clean. go.mod adds only hanzoai/notify + its provider transitive deps.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The fail-closed 503 body no longer names ZT_CLIENT_ID/ZT_CLIENT_SECRET;
it now reads 'networking is not configured on this deployment' — a
customer-facing string the console renders as a clean 'not available yet'
state. Ops still see the env names in the Warn log at Mount. No behavior
change; the gate() fail-closed contract is identical.
RED M-2 [MED] — contain panics in the async agent-turn goroutine. The dispatched
turn (handleSlack* → slackAgentReply → agents.RunOnBehalf, a large surface over
UNTRUSTED Slack input) ran UNRECOVERED — middleware.Recover() only wraps the sync
request goroutine — so a panic would crash the ENTIRE shared multi-tenant cloud
binary (every tenant, every subsystem). Introduced slackSpawn: runs an
already-slotted turn in a recovered goroutine (recover defer registered LAST so it
runs FIRST; the slot release still runs after it, so a panicking turn frees its
slot). Test TestSlackTurnPanicRecoveredAndSlotReleased.
RED M-1 [MED] — shed BEFORE burning the dedupe key. The old order was
mark-then-dispatch: MarkSlackEvent recorded the event_id, then a pool-full drop
silently 2xx-acked → Slack never retried and a later retry was deduped away, so
the @mention vanished. New order (events + slash): verify HMAC → resolve org →
TRY-ACQUIRE a pool slot; on shed record NOTHING and return a retriable 429 (Slack
re-delivers when a slot frees — the turn never ran, so no double-run and no burned
key); on acquire → MarkSlackEvent (release the slot on duplicate/error) → spawn.
The fast empty-2xx ack is kept for the normal path. Test
TestSlackShedReturnsNon2xxAndDoesNotRecord.
Also corrected the slack_dedupe.go replica note (we no longer "always 2xx-ack" —
a capacity shed returns a retriable non-2xx and records nothing, so it is not a
double-process).
RED M-3 [MED] deploy single-replica (Recreate + persistent CLOUD_DATA_DIR) — the
universe cloud manifest, owned by the deploy lane; the in-code single-writer
invariant comment is kept accurate here.
go build ./... = 0, go vet ./... = 0, go test ./clients/integrations/
./clients/agents/ -race green (25 tests).
RED H1 [HIGH] — wire the bridge into integrations.Mount (was only in a test
helper → the front-door didn't exist in prod). The 5 literal routes are
registered BEFORE the /:provider wildcards (registration-order precedence, same
discipline as clients/agents' static-before-:ref) and are PUBLIC at the JWT
layer: IdentityMiddleware only POPULATES a principal (never rejects) and
DefaultPrice returns 0 for /v1/integrations/* so BillingGate passes through —
reached exactly like /:provider/callback. Auth is HMAC (events/commands) /
signed __Host- cookie (link legs) INSIDE the handler. New test
TestSlackRoutePrecedence proves GET /slack/link hits slackLink (not :provider),
/v1/integrations/slack still resolves the provider view, and the webhook is
reachable with no principal.
RED M1 [MED] — per-org concurrency sub-limit. The agent-turn pool was one
process-global semaphore; one org bursting @hanzo could starve every tenant.
Replaced with orgLimiter (global cap + per-org cap, SLACK_AGENT_ORG_CONCURRENCY
default 8). The org is now resolved SYNC in the webhook path so the pool keys on
the RESOLVED tenant before a slot is taken. New test TestOrgLimiter.
RED M2 [MED] — corrected the false "holds across replicas" dedupe claim: the
table is per-process embedded SQLite (single-writer per HIP-0302), so the billed
webhook path MUST run single-replica (stated as the shipping invariant); a
shared SETNX store is the multi-replica follow-up.
RED L1 [LOW] — moved the dedupe-table DDL into the store's migrate() (store.go)
— fail-loud at Mount, one place — and removed the lazy first-use ensure whose
LoadOrStore-before-run could permanently disable the path on a transient DDL
error. slackBridgeReady now only inits the process pool + link seen-set.
Deferred (flagged for clients/integrations owner): L2 UNIQUE(provider,
external_id)+first-org-wins refusal on duplicate team connect; L3 purge
user:<slackUser>:refresh secrets on disconnect (currently inert after
disconnect, no leak).
go build ./... = 0, go vet ./... = 0, go test ./clients/integrations/
./clients/agents/ -race green (18 + 5 tests).
Port the hardened @hanzo Slack agent front-door from team-go/pkg/slack into
the unified Hanzo Cloud integrations plane, so Slack is ONE connector aligned
with the one-binary north star. It CONSUMES the existing Slack OAuth provider
(the per-org bot token it seals) and the framework seams
(OrgForExternalID / TokenFor / ConnectionFor); it adds no new custody path and
edits no existing file.
clients/agents:
- onbehalf.go: exported in-process RunOnBehalf(ctx, org, userSub, ref, input) —
the clean in-process twin of the HTTP run handler (no gateway hop, no
Cloudflare/IPv6 exposure). Resolves the agent org-scoped, runs it through the
SAME runAgent -> executeRun -> meter path, bills billingActor(org, userSub)
against org's ledger. Takes org+userSub DIRECTLY (caller pre-authenticated).
clients/integrations:
- slack_events.go: Slack Events webhook + slash command. HMAC-verified over the
EXACT raw body with a 5-min replay window; url_verification challenge; routes
@mention + DM to an on-behalf-of run; durable dedupe on event_id; fast empty
ack + bounded async worker pool. Posts the reply into the thread with the
org's bot token, or the link prompt EPHEMERALLY.
- slack_link.go: transplant-safe 3-leg per-user link (__Host- init/link cookies,
leg1<->leg2 nonce continuity checked BEFORE any exchange, single-use). Binds
Slack<->Hanzo via hanzo.id OIDC (hanzo-slack client) and seals the refresh
token per (org, "slack", "user:<slackUser>:refresh").
- slack_verify.go: Slack signature verify + single-use link-state crypto
(constant-time HMAC over s.stateKey; orthogonal to the OAuth-connect state).
- slack_dedupe.go: durable event-dedupe table as Store methods (no store.go edit).
PER-ORG ISOLATION (ship bar): an event's org comes ONLY from
OrgForExternalID(team_id) — never the payload; the reply uses THAT org's bot
token (TokenFor); the run is THAT org's agent (RunOnBehalf org-scoped). Tests
prove team A's event never resolves/tokens/runs as org B.
Mount wiring (5 routes) is handed to the clients/integrations owner — this
change adds NO Mount edit (clean separation); handlers are (s *svc) methods.
Tests (go test -race, green): HMAC reject (bad/missing/stale), dedupe
idempotency, per-org isolation (end-to-end bot-token capture proves the reply
used the connecting org's token), link transplant-rejected (no/mismatched init
cookie refused before exchange), RunOnBehalf bills the right actor.
The recorded zip h1 for github.com/luxfi/age v1.5.0 (zC/Fw…) did not match
the immutable Go checksum transparency log (sum.golang.org), which records
G69Hb… — the same bits the module proxy and local cache serve. The stale
hash made the ENTIRE module unbuildable: every `go build` failed with a
checksum mismatch / SECURITY ERROR. The /go.mod hash already matched sumdb;
only the zip h1 was wrong. Aligning it to the transparency-log-verified
value unblocks the repo (`go mod verify` -> all modules verified). age is an
indirect dependency; no version bump.
Cross-org fleet o11y for admin.hanzo.ai (global-admin only, s.guard fail-closed):
fleet totals (requests/tokens/cost/errors/orgs/models from hanzo.cloud_usage;
latency p50/p95/p99 + error-rate + services from signoz_traces; log volume from
signoz_logs), usage + log-volume timeseries, and top-N orgs/models/services
leaderboards, plus the fleet Langfuse generation rollup. Un-org-scoped by design
— the one place a fleet operator crosses tenants; a non-admin bearer is refused
403 before a row is read. Reuses the shared aiobject.DatastoreQuery transport
(no second connection) and the compute/analytics honest-empty pattern; admin
reads only, owns no table. Time bounds are positional params, bucket interval a
server-side constant — injection-safe. Pure builders + parsers unit-tested.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
luxfi/age@v1.5.0 was force-retagged upstream: proxy.golang.org now serves
content hashing to G69HbSV… while go.sum pinned the stale zC/Fw… → every
cloud release fails at `go mod download` with a SECURITY ERROR (checksum
mismatch), blocking the whole pipeline. Dockerfile already sets GOSUMDB=off,
so the mismatch is against the committed go.sum, not the sumdb. Realign to the
exact hash CI computes from the proxy (same fix IAM shipped as fee44857).
Addresses CTO's post-RED fix set (isolation boundary already approved airtight).
MED-1 (exactly-once metering/audit/persistence across ALL entrypoints): the
durable path is now the SINGLE owner of run bookkeeping. FlowRunWorkflow runs a
RecordRunStartActivity keyed on the workflow id (workflow.GetInfo — a scheduled
cron mints a fresh id per tick, so each tick is its own metered run despite the
schedule embedding one fixed FlowRunInput.RunID). Store CreateRunIfAbsent (row
idempotency) + ClaimMeter (atomic metered-flag 0->1) meter+audit only the winner,
so manual /run, MCP, and cron never double-bill. Manual /run no longer meters/
audits — it only CreateRunIfAbsent for immediate visibility. RecordRunEndActivity
records terminal status. Proof: TestScheduledRunMeteredExactlyOnce (a tick meters
once + shows in listRuns; two ticks = two distinct runs) + TestRunStartBookkeeping-
Idempotent (recordRunStart twice = one meter, one row).
MED-2 (honest SSRF blocklist): isPublicIP now rejects the IANA special-use ranges
Go's net helpers miss — 100.64/10 CGNAT (Alibaba metadata 100.100.100.200),
0/8, 192.0.0/24, 192.0.2/24, 192.88.99/24, 198.18/15, 198.51.100/24, 203.0.113/24,
240/4, 64:ff9b::/96 NAT64 — plus v4-mapped-v6 normalization. Comment no longer
overclaims a complete cloud-metadata blocklist. TestIsPublicIP covers each range +
public IPs still allowed.
MED-3 + LOW-4: step-count (<=256) + serialized-tree (<=512KB) caps at create /
version / operation time -> honest 422; resume payload bounded (<=64KB) -> 413.
LOW-2: per-org concurrency limiter (429) on run-starts + synchronous MCP tool calls
(bounds the core.delay goroutine lever). TestConcurrencyLimiter + TestFlowStepCap +
TestResumePayloadBounded.
LOW-1: MCP meters/audits AFTER Run, outcome derived from the real result — a failed
/ SSRF-blocked / not-connected call audits as error and is NOT billed. TestMCPAuditOutcome.
LOW-3: updateFlow validates publishedVersionId names an existing version OF THIS
FLOW in-org (else 422). TestUpdateFlowPublishedVersionValidated.
INF-1: register() panics at init on a <connector>_<action> tool-name collision so a
future connector can't silently make MCP dispatch ambiguous. TestToolNameCollisionPanics.
Tests: 25/25 green (CGO=0 build/vet/test; -race clean under cgo). Full module builds;
cmd/cloud links. catalog.json untouched.
Authorize needs only SLACK_CLIENT_ID (a public value in every consent URL);
the SECRET is required only at the callback token exchange. Gate available/
connect on client_id so an org reaches Slack's Allow screen as soon as the
public id is set, while a deployment still missing SLACK_CLIENT_SECRET fails
the exchange with an honest ?error=slack (never a dead-end).
Pulls the cors_filter static-allowlist fix so console.lux.cloud (and
zoo/pars brand consoles) stop getting 403 "origin is not allowed" on
/v1/signin. Cleared stale sum.golang.org-poisoned go.sum entries for
re-tagged luxfi/{age,precompile,keys} (GOPRIVATE direct re-records the
current content hashes; matches the repo's GOSUMDB-off CI recipe).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Native-Go /v1/automations/* subsystem in the unified cloud binary. Composes
three existing seams, never reinvents them:
- clients/integrations — per-org connector creds via integrations.TokenFor
(KMS-sealed, fail-closed); connectors never touch KMS directly.
- cloud.EmbeddedTasks — the ONE shared in-process durable engine; a flow
runs as a durable workflow in the OWNER's namespace (per-org lazy worker,
mirroring ai/object/ingest_tasks.go).
- clients/principal — the ONE tenant gate on every data handler.
Isolation is physical: ONE SQLite file, org column + org-led index on every
table; the durable activity's SOLE credential scope is FlowRunInput.Owner —
the VALIDATED org set at flow-start, never a client-supplied field.
Surface: pieces catalogue (go:embed), flows CRUD + versions + FlowOperation
apply, durable runs (start/list/get/resume via SignalWorkflow), enable/disable
(POLLING -> CreateSchedule), and a HIP-0300 MCP JSON-RPC tool surface
(/v1/automations/mcp) exposing every connector action as <connector>_<action>.
Connectors (Tier-A, self-registering): core (http_request SSRF-guarded via a
dialer Control hook, delay, code data-mapper, wait_for_approval signal
waitpoint), slack (send_message), github + google_sheets/drive (fail closed
until integrations custodies their tokens).
Metering + audit on flow-run start and MCP tool call. Order 148 (after
integrations 137, before ai's /v1/* catch-all 150).
Tests (16, all green with -race): store org-isolation, HTTP org-gating (403),
durable flow run reaching SUCCEEDED with threaded step outputs on an embedded
tasks engine, connector token isolation (in.Owner is the sole cred scope),
MCP tools/list + gated tools/call dispatch, pieces catalogue.
* feat(o11y): emit an OTel SERVER span per /v1/* request over the ZAP wire
Cloud installed a ZAP tracer provider (cmd/cloud initTelemetry) but nothing in
the handler chain opened a span, so no request ever flowed through it — the o11y
Monitoring tab saw zero hanzo-cloud request traces (receiver="zap" span count
was flat-zero while logs streamed over ZAP).
TracingMiddleware (middleware_tracing.go) opens one SERVER span per /v1/*
request off the GLOBAL tracer (= the ZAP provider), records the OTel HTTP
semantic-convention attributes (method, route, status) + request_id/org, maps
error/5xx to an error span status, and writes the span context back onto the
request via SetContext so every downstream span (agent.run -> agent.step -> the
chat client span in clients/aihttp) parents under it: one trace tree per
request. Health/readiness/metrics + non-/v1 paths are skipped so probes never
flood the trace store. Wired right after RequestID in the canonical pipeline
(serve.go) so the whole authenticated chain nests under it.
c.Path()/c.Method()/headers are zero-copy views over the fasthttp request
buffer, which is recycled for the next request BEFORE the batch span processor
serializes the span asynchronously — so retained views corrupt (live: a
GET /v1/models span exported with http.route="/v1/chat/c..."). strings.Clone
pins our own copy for every retained attribute. Tests cover emission, attribute
mapping, error status, parent/child propagation, the skip set, and an env-gated
on-wire live test (CLOUD_ZAP_LIVE_ENDPOINT) that ships real spans to a ZAP
receiver — the async-export + ctx-reuse path the in-memory recorder can't model
(and the one that surfaced the corruption).
* feat(o11y): make cloud the single tracer-provider owner — one wire (ZAP)
The fused cloud binary set the ZAP provider first, then ai.Bootstrap (during
MountAll) called hanzoai/ai object.InitTelemetry which, seeing the CR's
OTEL_EXPORTER_OTLP_ENDPOINT, installed a SECOND, competing OTLP provider. OTel
global delegation is first-writer-wins for handles created before the first
SetTracerProvider (cloud's package-level tracers keep ZAP), but the ai GenAI
tracer is resolved lazily AFTER the second Set, so its spans stranded on
OTLP(:4318) while ZAP owned the rest — the split that left receiver="zap" span
count at zero for hanzo-cloud (verified live: spans arrived only via
receiver="otlp").
Composition-root fix: once cloud installs the ZAP provider, clear the
OTLP-exporter env (OTEL_EXPORTER_OTLP_ENDPOINT / _TRACES_ENDPOINT) so no embedded
subsystem installs a competing OTLP provider. Exactly one provider (ZAP), one
wire, deterministic regardless of CR env drift. In the fused binary OTLP is only
ever the collector's interop RECEIVER, never cloud's exporter; standalone
cmd/aid (no ZAP endpoint) is unaffected and keeps its OTLP path.
---------
Co-authored-by: hanzo-dev <dev@hanzo.ai>
The auto engine now gates /v1/auto/pieces/{piece}/run on a shared secret (it
trusts X-Org-Id absolutely, so the write+SSRF surface needs an in-band caller
proof). cloud is the ONLY legitimate caller — it resolves each org's real token
and pins the provider URL — so it presents the secret (from KMS, the same
PIECES_RUNNER_SECRET) as X-Piece-Run-Secret. pieceSync fails closed if the
secret is unset (a doomed call the engine would 403 anyway).
Mounts workflow automation + the ~280-app activepieces long tail in-platform,
per-org, through the ONE knowledge store. Go core + JS on-demand.
clients/auto: /v1/auto/* per-org REVERSE PROXY to the standalone Hanzo Auto
engine (one engine, one store — not re-embedded). The auto engine trusts
X-Org-Id absolutely, so this proxy IS the trust boundary: it GATES on a
validated principal (refuses the anon-forge X-Org-Id-with-no-credential path)
and re-stamps outbound identity from validated values only (strips every
smuggled authority alias). Pure gate+proxy in clients/auto/proxy (5 isolation
tests: anon-forge 403, per-org forward, smuggled-header strip, path preserved).
clients/kb: the first LONG-TAIL connector (notion). Identical OAuth lifecycle
(HMAC-org-bound state, KMS token path) but its PULL runs the activepieces JS
piece through the auto engine's on-demand runner (sync_piece.go) instead of
native Go — then files each record via the SAME framework.Ingest path. One
ingestion path; a JS-sourced doc lands in the same per-org store+index as a
Go-sourced one. clients/kb/notion is the pure record-shaper (6 tests).
ONE catalog: /v1/kb/connectors/catalog lists native Go + long-tail piece
connectors in one list, each badged kind native|piece (3 tests).
RED LOW-1: collection() + kmsRef() now route org through provisioning.SanitizeOrg
(the codebase's ONE normalizer) so the physical Qdrant namespace + KMS path are
injective in the owner ("a b" != "a_b") — defense in depth under the payload.org
filter. Injectivity tests + KB integration tests updated to derive the collection
through the helper (robust to the normalizer).
All tests green under CGO=0 (production config). Full binary boots; /v1/auto
mounted, anon-forge 403, catalog gated, spine (kb/framework health) 200.
Address review: proxy the customer card list to commerce's admin-group PORTAL
endpoint (GET /v1/billing/portal/payment-methods), not the user-group
/payment-methods. PortalPaymentMethods 400s without a ?customerId=, so pinning
only ?user= would break it — generalize the proxy's subject pinning to the FULL
commerce edge-auth key set {user,userId,customerId} (now a shared
billingSubjectKeys var, identical to clients/console + commerce), pinned to the
caller's OWN org on every request. This leaves NO billing endpoint unfiltered
regardless of which param it reads (usage/balance/gpu-eligibility read user;
portal/payment-methods requires customerId) and is strictly more tenant-safe.
pinSubjectBody reuses the same var. The console keeps requesting the same-origin
/v1/billing/payment-methods (mounted here); the portal hop is server-side only.
Tests updated: widen-scope now asserts every subject key is pinned (org dropped);
payment-methods asserts the portal path + customerId pin. 14/14 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The customer billing proxy (clients/billing) exposed only usage+balance, so the
console GPU launch gate — card-on-file check + prepaid eligibility read + the
prepay-only charge — had no org-scoped cloud route and fell through to the
console pkg's /v1/billing/* wildcard (admin-shaped 403). Extend the SAME
commerce-proxy helper (no new HTTP/auth machinery) with the three enforcement
routes commerce 1.46.28 serves (api/billing/gpu_charge.go + portal):
GET /v1/billing/gpu-eligibility -> commerce GET (read-only launch gate:
{eligible,reason,prepaidAvailable,cardOnFile,...}; amountCents +
minPrepaidCents + currency pass through)
POST /v1/billing/gpu-charge -> commerce POST (prepay-only, card-required
debit; commerce enforces both gates + gpu-tagging server-side; status
forwarded verbatim: 201 ok / 402 card_required|insufficient_prepaid)
GET /v1/billing/payment-methods -> commerce GET (masked brand+last4 cards
for the card-on-file check; type passes through)
All org-scoped to the caller's OWN org from the VALIDATED IAM owner claim
(principal.Tenant), identical to usage/balance — a client can never widen scope:
the GET subject is pinned to ?user=<org> (commerce's privileged payment-methods
branch filters CustomerId on it), and the POST body subject is pinned to the
{user,userId,customerId} set (mirrors clients/console + commerce edge-auth), so
a forged body can never charge another tenant. New commerceProxy.post + the
pinSubjectBody helper; no principal -> 401, unconfigured -> 501.
Tests: gpu-eligibility scope+passthrough+forged-subject overwrite;
payment-methods scope+type; gpu-charge body-subject pin + 402 verbatim + 401
no-principal + 501 unconfigured; pinSubjectBody unit. 14/14 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
On console.hanzo.ai the ingress routes /v1/* straight to cloud-api:8000 (the
console Next BFF is only at "/"), so the console's /v1/billing/usage +
/v1/billing/balance calls land on cloud-api — NOT the console's per-tenant
commerce proxy. cloud-api wired commerce billing ONLY under the admin-gated
aggregate (/v1/admin/*), so a normal org owner (davelorenzini/maxpower) hitting
/v1/billing/usage had no customer route and was denied -> 403 -> the "Access
required" wall on EVERY product overview + o11y usage panel.
Add a customer-facing, org-scoped billing READ surface (clients/billing):
GET /v1/billing/usage and /v1/billing/balance. Org = the VALIDATED IAM owner
claim (principal.Tenant — the trusted X-Org-Id the identity middleware minted
from the caller's verified session; never a client header), so a customer reads
ONLY their OWN org. Proxies commerce with COMMERCE_SERVICE_TOKEN + X-Org-Id=<org>
and the per-org billing subject pinned to user=<org> (admin.orgSubject /
metering identityFromCtx — verified live: user=<org> returns the real wallet);
returns commerce's raw body + status verbatim (the console parses the raw ledger).
Tenant isolation: no client-supplied subject/org query is ever forwarded, so
scope can never be widened. The all-orgs god view stays admin-only (clients/admin).
Tests: subject pinned to caller org, forged user/userId/customerId/org dropped,
start/end/currency pass through, no-principal -> 401 (commerce untouched),
unconfigured -> 501, commerce status forwarded verbatim.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pulls ai#68: async OpenAI Sora-style video API onto the cloud router.
POST /v1/videos/generations returns a video_<uuid> job immediately (was
sync ~104s → console /ai proxy 502); GET /v1/videos/{id} polls; GET
/v1/videos/{id}/content streams the MP4. Metering exactly-once (hold on
create, settle on completion, reaper releases abandoned), ownership-secured.
ai v1.800.0 go.mod is identical to v1.799.3 — no dependency-graph change;
only the hanzoai/ai hash lines move. Verified: cloud binary builds clean
(CGO_ENABLED=0) and the router now serves /v1/videos/generations,
/v1/videos/:id, /v1/videos/:id/content (spark-video backend).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Durable replacement for the Huly/Svelte hanzo.team tracker whose upstream
each-block reactive-batching render race left issue lists rendering zero
rows. Native Go over one org-scoped SQLite store: rows return as plain JSON
and render deterministically (no Svelte reactivity in the path).
clients/tracker/store.go — projects + issues on {DataDir}/tracker.db, the
same modernc/SQLCipher driver + MaxOpenConns(1)+WAL pattern as projectsvc/
crm. Per-project monotonic issue numbering allocated inside one tx (never
races under the single-writer conn). Cascade delete in a tx. org column is
the tenancy key; every query filters WHERE org=?.
clients/tracker/tracker.go — /v1/tracker/projects[/:key][/issues[/:num]] CRUD.
org = principal.Tenant (validated IAM owner claim, HIP-0026), 403 otherwise.
Status/priority closed sets; board/list via ?status=. Create wired to the
shared per-org billing seam (free by default; ops prices via
CLOUD_TRACKER_FEE_CENTS). Registered order 129, before the AI /v1/* catch-all.
subsystems/subsystems.go — one blank import links it into the binary.
Store CRUD/numbering/status-filter/cascade/tenant-isolation proven green on
real SQLite (clients/tracker/tracker_test.go).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
ROOT CAUSE of durable ingest silently running inline (round-trip: github ingest blocked
25s+, no workflow in Tasks): the embedded engine only registers 'default' at boot and does
NOT lazily create namespaces on ExecuteWorkflow, so dialing Namespace:<org> made the worker
poll a non-existent namespace and BLOCK → EnqueueIngest hung → handler fell back to inline.
Fix: dial 'default' (always registered). Data isolation is unchanged — it's in the workflow
INPUT (IngestSource is owner-scoped), never the namespace. Now github/crawl enqueue a real
durable workflow that appears in Tasks (under default).
V6: accept the owner-bound per-tenant machine audience (<owner>-platform-kms) so a real client_credentials sync token clears SanitizeIdentity + the /v1/kms guard; decoupled from global-admin (isKMSMachinePrincipal). V1: per-IP rate limit + MaxConnsPerHost(h2-off) on the public login broker. Blue→Red→Blue→Red: Red SHIP (0 crit/high/medium). Activation runbook in the PR body + EnsureOrgIdentity doc.
Fourth app lane (after cms/erp/help): a Notion-like KB + agent memory + app
connectors as a 'kb' module on clients/framework — no new Base, no new database.
Fixtures (module kb): kb-page (wiki tree via a self-Link parent + Lexical
RichText body), kb-memory (agent memory: note/fact/observation), kb-source
(connector-ingested docs), kb-connector (connection metadata; OAuth token in KMS,
never in the doc/logs). All CRUD/permissions/tenant-isolation/install are the
framework's generic /v1/framework/* surface + the generic @hanzo/ui renderer.
Indexing (index.go, the ONE vector-write path): an after_save hook embeds every
knowledge write (page/memory/source) via the gateway and upserts it into the org's
OWN Qdrant collection (kb_<org>) with an org-pinned payload; on_trash removes it.
Human wiki + AI memory are ONE per-org knowledge store, indexed once. Fail-open at
index time (a vector outage never blocks a knowledge write); fail-honest at query.
Retrieval (subsystem.go): POST /v1/kb/search is the org-scoped RAG entry point —
collection AND payload filter both pinned to principal.Tenant, so a caller can only
retrieve its OWN knowledge. Degrades to an honest empty result when the index is
down.
Connectors (connectors.go, sync.go): per-org OAuth to GitHub/Slack/Google that
ingest external docs INTO the same store + index (via framework.Ingest → same
after_save hook — one ingestion path, never forked). OAuth state is HMAC-bound to
the validated org (defeats login-CSRF/mix-up); tokens live in KMS at a per-org path.
GitHub is end-to-end (repo READMEs + issues); Slack/Google share the OAuth
lifecycle + normalizer with an honest 'listing not yet implemented' depth marker.
framework.Ingest/UpdateData/FindByField/Search: the in-process create-with-hooks
API first-party producers (the connector sync) use so off-request writes run the
exact validate + lifecycle pipeline and stay physically org-scoped.
Tests (16, all green): engine-Validate on every fixture; page self-Link + Lexical
body; connector-has-no-token; per-org pointID/collection/kmsRef isolation; payload
org-pin; OAuth-state org-binding + tamper/cross-provider/wrong-key rejection;
Ingest validates+fires-hooks+org-scoped. Integration (real SQLite + mock
Qdrant/embeddings over the real HTTP surface): install → create kb-page in org A →
after_save indexes into kb_A → search as A retrieves it → search as B sees NOTHING
(no cross-org leak); forged-org search → 403; kb-memory lands in the same kb_A
namespace.
Adversarial review of the connector framework (state HMAC, nonce custody,
per-org KMS token custody, console redirect). Core design held: state forgery,
nonce replay/race, cross-tenant KMS pathing, principal-forge, and the seam
fail-closed contract were already sound. Fixes are defense-in-depth + red-style
proving tests; the contract is unchanged (no /api/, no /v1/slack route).
Fixes
- Ingest sanitization (vector 7): provider-supplied NON-secret metadata
(account label / external id / bot user id / scopes) is now stripped of C0
control chars + DEL and length-bounded at the ONE framework ingest point in
callback, before it is logged, stored, or reflected. Kills log-line/separator
injection via a crafted Slack workspace name and bounds per-org row growth.
Secret token VALUES bypass this and go straight to the KMS seal.
- Open-redirect hardening (vector 4): success/failRedirect fold into one
query-escaping consoleRedirectURL builder (DRY + unit-testable). Confirms the
Location host is always the env-fixed console origin; hostile provider detail
can't break out of the query into host/scheme/path or inject CRLF.
- Request bounds (vector 10): callback rejects an oversized OAuth `code`
(maxCodeLen, also covers the in-process ZAP plane); verify() rejects an
oversized state token before any base64 work (maxStateLen).
- kmsDelete uses errors.Is(kms.ErrSecretNotFound) for wrap-safe idempotency.
Tests (all real, -race green; 39 pass)
- state: dot-injection/degenerate split, MAC-checked-before-parse,
validly-signed-but-hostile-org rejected, overlong rejected.
- store: concurrent single-winner nonce consume (race proof).
- http: no-open-redirect property, end-to-end metadata sanitization,
disconnect anonymous-forge -> 403 (secret+row survive), github scaffold
callback fails closed at the Configured gate before any exchange,
oversized-code rejected.
Provider-agnostic /v1/integrations plane: one registry, N providers.
Slack = full reference impl (OAuth v2 bot-token); GitHub = scaffold + #51 seam.
Per-org token custody in KMS (sealed); state-authed HMAC callback with
single-use nonce; org derived ONLY from signed state on the public callback.
Generic /v1/integrations/{provider}/callback — no /v1/slack/* (team-go owns it).
Wired at order 137 (after security 136, before AI 150).
release.yml — invert the tag/build order so a git tag can NEVER exist without a
pushed, boot-verified image (the phantom v1.786.42/43 → ImagePullBackOff cause):
main push → compute next version → build → SMOKE (boot to "listening") →
push image → git tag (receipt) → notify universe
- Tag is minted only AFTER the push step succeeds; any build/smoke/push failure
fails the run before the tag step → fail-run-no-tag.
- concurrency group `release-cloud` (cancel-in-progress:false) serializes runs so
two main pushes can't collide on a number; the queued run re-reads tags and
lands on the next patch → monotonic.
- next version = max(highest git tag, highest pushed ghcr container tag) + 1,
folding in container tags so a pushed-but-untagged number is never reused.
- removed the `tags: v*` trigger (this workflow now OWNS tags — a hand-cut tag has
no image behind it and won't build); notify-universe fires only on a successful
build+tag, so universe is never told about a phantom.
sqlite — fix `panic: sql: Register called twice for driver sqlite` under
CGO_ENABLED=1 (blocks `go test ./...` + clean cgo rebuilds). #96 moved cloud's
stores to github.com/hanzoai/sqlite (mattn under cgo) while several embedded deps
still import modernc.org/sqlite directly (ai, tasks, base, commerce, o11y, orm) →
two packages register "sqlite" under cgo. Prod is CGO_ENABLED=0 (all one modernc
package, deduped) so prod never paniced; the panic is cgo-only.
- bump github.com/hanzoai/sqlite v0.1.4 → v0.1.5: adds the `sqlite_purego` opt-out
build tag that forces the fork's pure-Go (modernc) backend under cgo. Default
cgo path is unchanged (mattn/SQLCipher) so IAM/commerce encryption is untouched.
- bump github.com/hanzoai/ai → the commit that routes object/adapter.go + cmd
tools through hanzoai/sqlite instead of modernc (never modernc directly).
- Makefile: CGO_ENABLED?=0 default (matches the shipped Dockerfile) so `make
build`/`make test` register "sqlite" once and exactly mirror prod; new `test-cgo`
target proves the cgo path via `-tags sqlite_purego`.
Verified: CGO_ENABLED=0 `go build/test ./...` and CGO_ENABLED=1 `-tags
sqlite_purego go build/test ./...` both pass with NO panic (the eval package that
panicked now passes in both modes). Pre-existing clients/s3 + clients/functions
billing-attribution test failures are unrelated (present on clean main, both
modes) and out of scope.
billing.go's isSafeSegment left percent-escape (`%2f`/`%2e`) and matrix-param
(`;`) segments undecoded, so `/v1/billing/x/..%2fadmin` forwarded
`x/..%2fadmin` verbatim; the Go http client + commerce's own router
decode+normalize it downstream into a path that tunnels PAST /v1/billing into
another surface. commerce.go had already patched its call site with an inline
`%;` check — braiding the policy across call sites.
Harden the ONE shared segment guard instead: isSafeSegment now rejects empty,
`.`/`..`, slash, backslash, percent-escape, matrix-param, and any control char.
Both bridges (billing + commerce) get the complete guard from one place, and
commerce.go's call site drops the now-redundant inline check.
Regression test: `..`, `%2f`, `%2e%2e`, and `;` all 400 and never reach
upstream (proven to fail against the pre-fix guard).
The store twin of the just-merged /v1/billing/* bridge (#102). #81 namespaced
the console's commerce store calls to the canonical same-origin /v1/commerce/*
(SPA->server->/commerce proxy); the statically-exported console now terminates
every dynamic call at the unified cloud binary's /v1, so the binary must
reverse-proxy /v1/commerce/* to the commerce service. Without it the commerce
embed was incomplete.
clients/console/commerce.go serves GET|POST|PUT|PATCH|DELETE /v1/commerce/<path>
-> commerce's BARE store surface /v1/<path> (the console-side 'commerce'
namespace is stripped: the deployed commerce cmd/commerced mounts
api.Route(Group('/v1')), so products/orders/customers/... live at /v1/<kind>
while money lives at /v1/billing/*). Exactly the mapping console2's next.config
rewrite proved live (/v1/commerce/:path* -> /commerce/v1/:path* ->
commerce.svc/v1/:path*).
IDOR-safe: the org is the VALIDATED caller's own (resolveCaller ->
principal.Validated / c.Org()), never a client value; a bearer-less forged
X-Org-Id has no validated principal and is refused 403 before any commerce call.
Reuses the commerceDo(base,token) S2S transport billing.go/topup.go share
(admin COMMERCE_SERVICE_TOKEN + X-Org-Id, which commerce's EdgeAuth trusts only
behind the service token). Least privilege: a store-head allow-list (identical
to console2 proxy-allow.ts COMMERCE_HEADS) so the bridge can never tunnel to
/v1/billing (its own subject-scoped bridge), /v1/checkout, or tenant admin.
Hardened over a naive port: rejects percent-encoded path segments (%2f/%2e),
which the router leaves undecoded in the wildcard param but the Go http client +
commerce's router normalize downstream -- 'product/..%2fbilling' would otherwise
tunnel to /v1/billing past the allow-list (RED). Mirrors console2 pathIsClean.
/v1 only. CGO_ENABLED=0 go build ./... ok; go test ./clients/console/ ok.
Companion to #89 (KMS-sealed PaaS secret env) + universe #321. The sync was
INERT for two structural reasons this closes, and the per-tenant scoping the
task requires is now enforced at cloud's ONE auth boundary — proven by test.
WHY IT WAS INERT (coordinate drift + wrong CR shape):
- Seal path ≠ read path. cloud sealed at /platform/tenant-<org>/<app> but the
kms-operator reads through cloud's org-scoped surface /v1/kms/orgs/<org>/
secrets/... which folds to /orgs/<org>/... — a DIFFERENT record, never found.
- The CR set projectSlug="platform" (a literal) and omitted secretsScope.keys,
which the CRD REQUIRES (MinItems=1; luxfi/kms has no list endpoint). Either
alone starves the sync.
- hostAPI carried a /v1/kms suffix; the operator appends /v1/kms/... itself, so
login + read URLs doubled the prefix.
THE FIX (secrets.go):
- Seal at the org-scoped coordinate orgs/<org>/platform/<app>/<KEY> — the EXACT
store path cloud's org-scoped read surface addresses. Seal and read are now one
coordinate (proven: TestPaaSSecretSealReadAlignment).
- CR carries projectSlug=<org>, secretsPath=platform/<app>, envSlug=default, and
the explicit sorted key roster. hostAPI = KMS root.
PER-TENANT SCOPING (the NON-NEGOTIABLE) — enforced, not hoped:
- The operator authenticates as a per-tenant IAM machine identity (owner=<org>)
via the NEW /v1/kms/auth/login broker (kmssvc/login.go): it exchanges the
caller's clientId/clientSecret at IAM's client_credentials endpoint and returns
IAM's owner-scoped token verbatim. cloud is a relay, not an issuer.
- cloud's org-scope guard admits /orgs/<org>/... ONLY when the VALIDATED owner ==
that org (SanitizeIdentity derives owner from the token, ignoring client
X-Org-Id). So tenant-A's credential can NEVER read tenant-B's path — 403 before
the store is touched (proven: TestPaaSSecretCrossTenantDenied).
- credsSecret is a PER-TENANT name in the tenant namespace — never a shared
platform-wide reader (that would be a cross-tenant hole = NO-SHIP).
PROVISIONING (ensureTenantKMSAuth) — fail-closed, one privileged seam:
- On app-create/deploy cloud ensures the tenant's owner=<org> credential is
projected into tenant-<org> as the creds Secret the CR references, via an
injected tenantKMSIdentity provider. nil provider (default) ⇒ honest "pending"
(operator can't log in ⇒ reads nothing) — NEVER a shared or wrong-org identity.
- FLAGGED: the concrete provider needs a scoped IAM admin credential cloud does
not yet hold (clients/admin/iam.go replays the caller's cred, no service
identity). Until wired/verified, sync stays safely pending. Flip ON = provision
the per-tenant identity; the security invariant holds regardless of how it is
minted (the guard is the enforcement point).
Tests (CGO_ENABLED=0 — prod/CI build mode; CGO test builds double-register sqlite,
a pre-existing module issue): alignment round-trip, cross-tenant 403 (+ A→A 200,
B→B 404, unauth 403), login broker happy/bad-cred/malformed, per-tenant provisioning
(fail-closed when unprovisioned, org-bound projection, no cross-tenant ask). Existing
kmssvc red-team guard vectors unchanged and passing.
The BFF catch-all sweep's one real "server work -> Go handler" case. console2's
app/billing/v1/[...path]/route.ts injected the commerce SERVICE token and pinned the
caller's billing subject server-side (work a static export cannot do). Ported to
clients/console/billing.go: GET|POST /v1/billing/* forwards to commerce with the admin
COMMERCE_SERVICE_TOKEN, scoping every request to the VALIDATED caller's own subject
(billingSubject + scopedBillingSearch + scopedBillingBody, the Go port of console2's
billing-scope.ts), so a tenant can only read/act on its OWN ledger.
IDOR-safe: the subject is the validated principal (resolveCaller: principal.Validated /
c.Org() / c.User()), never a client userId/org. A forged X-Org-Id with no validated
X-User-Id is refused (403). Unset COMMERCE_SERVICE_TOKEN -> honest 501.
DRY: commerceDo (topup.go) refactored to take (base, token) so the wallet top-up AND
this bridge share one S2S transport. No behavior change to topup (its tests pass).
Tests (billing_test.go): billingSubject personal/dedicated, subject pin + org drop +
passthrough (query & write body), forged-value overwrite, 403 no-principal, 501 no-token,
and the end-to-end scoped forward to a fake commerce. go build ./clients/console/ = 0;
go test (CGO_ENABLED=0) ok. The default-CGO modernc-vs-CGO sqlite double-register that
panics the package is a pre-existing repo-wide issue (separate sqlite-one-driver lane).
Co-authored-by: Hanzo AI <ai@hanzo.ai>
SanitizeIdentity minted an un-forgeable X-Org-Id but passed the org sub-scopes
X-Project-Id / X-App-Id through verbatim, so a caller could assert ANOTHER org's
project as a compute_usage attribution key or per-project sub-scope. Sanitize the
sub-scopes in the ONE trust boundary:
- Delete every X-Project-Id/X-App-Id on ingress (no raw client copy survives),
then re-inject only for a validated principal, against the acted-as org.
- Refuse a cross-org X-Project-Id: a project REGISTERED to a DIFFERENT org than
the validated org is dropped; the caller's own registered project and
unregistered free-form within-org labels survive (projectIsForeign; fail-closed
on a registry error).
- Drop both sub-scopes entirely on the anonymous path.
- X-App-Id is a caller label, not an isolation boundary (no cloud subsystem
scopes access by it; the un-forgeable org bounds any mislabel) - forwarded on
the validated path, dropped when anonymous.
Dependency-inverted like sites.SetResolver: projectsvc registers a
TenantScopeResolver at Mount; cloud never imports the project registries. The
visor proxy forwards the now-validated sub-scopes so compute attribution lands on
the caller's own project. X-Org-Id anti-forgery is unchanged.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Add a RichText fieldtype to the DocType engine so a content field can hold a
Lexical EditorState JSON string (the console renders it with a native WYSIWYG).
The value is opaque text — validate.go coerces it verbatim (clipped to the scalar
bound), stored in the schemaless doc blob, round-tripping through create→get with
no shape enforcement (that's a UI concern). Minimal + fail-closed: the fieldtype
is added to the validated allow-set, so an unknown type is still rejected.
CMS: the seeded content body (Article/Page/Post) becomes RichText, and a
Data field is added for optional per-project scoping (the console's org→project
switcher filters content by it; empty = org-level). One CMS engine; project is a
filter.
Tests: TestDocTypeValidate accepts a RichText field; TestFieldTypeValidation
proves a Lexical JSON string round-trips verbatim through validation.
The ML control plane (clients/ml, /v1/ml) ran under the pod SA (cloud-api),
braiding KServe/Kubeflow cluster reach onto the product-API identity. newDynamic
now re-scopes the in-cluster client to a mounted cloud-ml ServiceAccount token
when HANZO_ML_TOKEN_FILE is set (keeps in-cluster host+CA, swaps only identity),
fail-closed if the configured token is unreadable. Unset -> unchanged behaviour
(pod SA) for local/dev and pre-cutover. Pairs with universe ml-rbac.yaml
(cloud-mlsvc ClusterRoleBinding -> cloud-ml). go build + go vet clean.
The execution queue for "all Hanzo Go services merge into the one cloud binary"
(HIP-0106): wave 0 already-merged (8 embedded modules + 37 native clients),
wave 1 tasks+visor (this build), wave 2+ mount queue (notify2, extract-svc,
playground, ...), and keep-standalone with reasons (iam/gateway/kms-MPC/registry
/s3/docdb/chain daemons). Waves are sequential through go.mod to avoid the
in-flight collision this wave hit.
Consolidates the Tasks product surface into cloud — the follow-up durable.go
named ("consolidating that surface into cloud is the follow-up"). durable.go
already embeds the ONE tasks engine in-process (loopback ZAP :19999) for ai's
durable ingest; this mounts THAT SAME engine's HTTP handlers, so the Tasks
UI/API reads the same durable state as ingest. One engine, one binary, one way —
no second Embed.
- clients/tasksvc (order 147, before ai's /v1/* catch-all): adapts the shared
engine's HTTP surface onto the zip mux at /v1/tasks/* + the embedded React UI
at /_/tasks/*. The engine is created after MountAll, so the surface resolves
cloud.EmbeddedTasks() lazily per request (503 fail-soft until wired).
- gate: settings/cluster/health stay open (no per-org data); the data surface
(namespaces/workflows/mcp/events) refuses an unvalidated principal (403, never
the unscoped store) and threads the gateway-validated org into the engine via
tasks/pkg/auth.WithIdentity — per-(org,ns) shard isolation, matching the rest
of the cloud data plane (clients/principal).
- durable.go: export EmbeddedTasks() — the single shared-engine accessor.
- bump hanzoai/tasks v1.43.0 → v1.46.0 (in-proc identity seam + one sqlite
driver). No local replace directives.
Proof: /v1/tasks/cluster returns nodeId "cloud-tasks" (durable.go's engine),
settings 200, data routes 403 without a principal, /_/tasks UI 200.
ERP (clients/erp): ERPNext-core DocType fixtures + native-Go GL/stock hooks —
idempotent deterministic-leg postings (exactly-once under concurrent submit),
on_cancel reversal, finite-guarded totals, double-entry submit gates.
Help (clients/help): Frappe Helpdesk-core fixtures, pure DocTypes, no hooks.
Both register on the framework engine at init; installed per-org via
/v1/framework/modules/{erp,help}/install. No new HTTP surface — ERP/Help ARE
documents on /v1/framework/*, drawn by the same generic DocType renderer as CMS.
Verified CGO=0 (production Dockerfile config): binary boots clean (no double
sqlite driver register), /v1/framework/health 200, /v1/models unshadowed (503
real AI handler), erp+help tests green under -race.
Mounts /v1/bots in cloud as the sibling of /v1/machines — a Bot is an
Agent(cloud /v1/agents) + a kind=bot Machine(vm) + their AgentBinding,
composed as a thin proxy over the SAME Visor client the machines routes use:
GET /v1/bots list (vm /v1/machines?kind=bot + bindings join)
POST /v1/bots/launch machine launch{kind:bot} THEN bind-agent
GET /v1/bots/:id machine + its binding (404 if not a bot)
DELETE /v1/bots/:id unbind THEN terminate the machine
POST /v1/bots/:id/:action message=agent run | stop|pause=unbind
Plus the machine agent-binding proxies cloud lacked (vm already serves them):
POST /v1/machines/:id/bind-agent
GET /v1/machines/:id/agent-binding
DELETE /v1/machines/:id/agent-binding
GET /v1/agent-bindings
Every route org-gated by the validated principal (principal.Tenant), forwarded
to vm as ?owner=<org> — 403 without a valid IAM owner, exactly like machines.
No vm change: kind=bot launch + bind-agent are already live at visor:19000.
message runs the bot's bound agent via the ONE agent runner (/v1/agents/:agent/run).
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Embed's default os.MkdirTemp("") resolves to /tmp, absent in the distroless cloud image
→ 'tasks.Embed: tempdir: stat /tmp: no such file or directory' → fail-soft to inline.
Pin DataDir to {deps.DataDir}/tasks (MkdirAll first) so the in-process engine actually
boots. Verified: without this the warn+inline-fallback fired cleanly in prod (v1.786.48).
cloud embeds hanzoai/tasks IN-PROCESS (loopback ZAP, durable.go) and injects a per-org
dialer into ai's ingest — there is no external tasks service to auth to, no per-org token
minting, no HTTP inner-cloud hop. Long ingests (github/crawl/s3) run as durable workflows
in the owner's namespace (CONTRACT §6); upload stays inline. Fail-soft: embed error →
ai dialer unset → inline fallback. Bumps ai → v1.796.4 (per-org ingest dialer). One engine,
one binary, one way. NOTE: embedded store is memdb today (survives worker crash via retry,
not process restart); console /tasksd still points at the cluster tasks Service for the UI
(consolidating that surface into cloud is the follow-up).
Red verdict FIX-THEN-SHIP (0 critical; all findings within-tenant integrity —
isolation, forge-proofing, gates, ledger perms, console deletions all refuted/solid).
- HIGH (TOCTOU double-post): on_submit postings ran before the atomic docstatus flip
with hash-named legs, so N concurrent submits over-posted the ledger 4-6x. Every
GL/stock leg now has a DETERMINISTIC name (voucher-<kind>-<index>, via prompt
autoname) and postLeg is idempotent (pre-read + re-check on create-conflict), so
posting is exactly-once under any concurrency AND replayable after a partial
failure — no engine change, uses the existing store API. Balances stay SUM(ledger).
- MED (cancel did not reverse): on_cancel hooks append reversing ledger rows (swap
debit/credit; negate qty), sharing the SAME leg computation as submit.
- LOW (non-finite total -> 500): finite guards in the totals hooks -> clean 422.
- LOW (comment): tightened the ledger-immutability doc — manager bypass is
within-tenant authority only (Red confirmed no cross-org escalation).
Regressions: 8 concurrent submits -> exactly 1x200 + GL posted once (2 legs,
race-clean); cancel -> net GL zero-sum + net stock zero; overflow qty*rate -> 422.
go test -race 9/9 (CGO=1) + CGO=0 + vet clean + full cmd/cloud binary.
Every AI call already writes the proven hanzo.cloud_usage ledger (model,
provider, tokens, cost_cents, org, user, status) via ai/object's zapWriteUsage
— the same recordTrace funnel that emits the OTel GenAI span. This mounts the
missing native route the console Observe > Observations surface calls, reading
that ledger as Langfuse-v3 GENERATION observations (org-scoped by the validated
principal, bound positional params, bounded LIMIT). No new emission path: one
recordTrace, fanned to o11y (span) + Langfuse (span) + this ledger (read).
- telemetry.go: Observation model + ObservationFilter; ListObservations on the
Telemetry interface, dsTelemetry (cloud_usage query), memTelemetry (honest
empty); asInt64 coercer (cloud_usage UInt32 tokens / UInt64 cost — asFloat
only handles float types).
- eval.go: listObservations handler + observationView/toObservationView mapping
to the console Observation shape; GET /v1/evals/observations route.
- observations_test.go: view mapping (success/error), asInt64 coercion, mem
telemetry empty + org-required.
Requires a cloud rebuild + deploy-by-sha to go live (route is code, not env).
Co-authored-by: hanzo <a@hanzo.ai>
- zaptrace: otlptrace.Client over the ZAP wire (github.com/zap-proto/http) —
spans marshaled as OTLP protobuf, shipped over ZAP frames, NEVER OTLP-HTTP
(:4318)/gRPC(:4317). Target = collector zapreceiver (:4319).
- cmd/cloud/telemetry.go: initTelemetry uses the ZAP exporter; enable via
OTEL_EXPORTER_ZAP_ENDPOINT (default otel-collector.hanzo.svc:4319); keeps the
no-op-when-unset posture.
- clients/aihttp.go ChatCompletion: OTel GenAI client span (gen_ai.system/
operation.name/request.model + response.model + usage.{input,output}_tokens),
RecordError on failure. Captures the previously-discarded resp.Usage.
- clients/agents/agents.go: runAgent opens a per-run root span (agent.run),
executeRun a child agent.step span; the LLM client span nests under them —
one trace per run: run -> step -> chat.
Test: zaptrace TestUploadTracesOverZAP green (span over the real ZAP wire).
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Supersede the Base swap (67bb0bc). Hanzo Base has a REST/realtime document
API but NOT the MongoDB wire protocol, so it is not a drop-in docdb — a
customer's mongodb:// driver cannot speak to it. The managed "document
database" must accept existing MongoDB drivers unchanged.
The per-org dedicated docdb instance is now a FerretDB v1.24 instance
speaking the MongoDB wire protocol on :27017, backed by Hanzo SQL (SQLite,
pure-Go) — Mongo databases map to SQLite files under /state, collections to
tables, documents to JSON1 rows. ZERO raw mongod (no WiredTiger), ZERO
Postgres, ZERO go.mongodb.org driver in cloud (FerretDB is a deployed pod,
not a Go import). Per-instance SCRAM auth via FerretDB's SQLite-backend
new-auth (FERRETDB_TEST_ENABLE_NEW_AUTH + FERRETDB_SETUP_*); the returned
password is the instance admin credential, sealed in KMS.
engine fields:
- image ghcr.io/hanzoai/docdb-sqlite:1.24.0 — pinned FerretDB v1 with the
SQLite backend handler, mirrored from upstream by hanzoai/docdb CI (v2 and
the hanzoai/docdb Postgres/DocumentDB fork both dropped SQLite; v1 is the
last line that carries it). Distinct package from the Postgres-backed
ghcr.io/hanzoai/docdb that backs shared chat-docdb.
- fsGroup 1000: FerretDB is distroless and runs as UID:GID 1000 (no
entrypoint can chown), so a fresh block PVC must be group-writable via the
pod securityContext.fsGroup the operator stamps from spec.fsGroup — else
the instance CrashLoops on "permission denied" writing /state.
- FERRETDB_STATE_DIR + FERRETDB_SQLITE_URL pin both process state and the
SQLite files onto the mounted /state PVC (persist across restarts).
Verified end-to-end against the FerretDB v1.24 SQLite image with this exact
env: mongosh Insert/Find/Update/Delete over the wire protocol, and /state
held per-db admin.sqlite + events.sqlite with "SQLite format 3" magic — no
WiredTiger datadir, no Postgres PG_VERSION (both locally and in hanzoai/docdb
CI on the mirrored image).
TestDedicated_DocdbIsFerretOnSQL asserts the FerretDB image, SQLite backend
env, mongodb:// connString, per-instance SCRAM credential, fsGroup 1000, and
the absence of any Postgres/IAM/Base env. unavailableKinds stays empty.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Second and third app lanes on clients/framework, reusing the generic DocType
engine + install path + generic renderer (like clients/cms) — zero forked
engine, zero HTTP surface of their own, per-org on Base/SQLite.
- clients/erp: 20 ERPNext-core DocTypes (module "erp") — masters (item/
warehouse/customer/supplier/account/department/employee), submittable
transactions with child Tables (sales-order/-invoice/purchase-order/
stock-entry/journal-entry/payment-entry), and two read-only hook-posted
ledgers (gl-entry, stock-ledger-entry). Business logic as native-Go hooks:
line/document totals (before_save), balanced GL on invoice/journal/payment
submit, append-only stock ledger on stock-entry submit, and double-entry /
non-empty submit gates. Posting is org-scoped via ev.Org + ev.Store.
- clients/help: 5 Helpdesk DocTypes (module "help") — hd-ticket (status
workflow) + hd-agent/hd-team/hd-sla/hd-canned-response. Pure fixtures, no
hooks — the purest DRY proof; self-contained (no cross-lane Link).
- SLUG names (erp-*/hd-*), series naming for transactions, field naming for
masters (console slugifies on write), hash for ledgers — all reachable via
the generic renderer; no collision with CMS (Author/Media/Page/...) or CRM.
- subsystems: blank-import erp + help so their init() registers the lanes.
Tests: fixtures-valid/model-spec/install-transact-roundtrip/submit-gates/
tenant-isolation/ledger-read-only — go test -race green (CGO=1 and CGO=0),
go vet clean, full cmd/cloud binary builds.
The pure detection engine moves to clients/security/detect (a stdlib-only
LEAF): one engine, two surfaces — the /v1/security HTTP subsystem and the
new local CLI both consume it, neither drags the other in. The subsystem
now calls detect.ScanContent/Rules/SeverityRank; behavior is unchanged.
hanzo security scan [path...] walks a tree, runs the engine, and exits
non-zero when a finding at/above --fail-on (default low; 'none' = report
only) is present — a pre-commit/CI/agent guardrail with no server, auth,
or network. Skips vendored/binary files; never prints a raw secret (masked
preview only). hanzo security rules lists the catalog. -o json supported.
Tests: 9 CLI (find+fail, clean-pass, fail-on threshold/none, json, vendor+
binary skip, bad flag, rules, control-verb), engine+subsystem unchanged.
Extends the white-label issuer-set validation (this branch) to the AUDIENCE half:
a lux/zoo/pars session token carries aud=<brand>-cloud (HIP-0111: client_id == app
== aud), so the audience allowlist must include each or the cloud-native identity
sanitizer 401s a valid lux token even after the issuer gate passes.
- brand.go BrandAudiences() derives <brand>-cloud for every registry brand
(hanzo-cloud, lux-cloud, zoo-cloud, pars-cloud, bootnode-cloud) — one source of
truth, mirroring BrandIssuers(); no hand-listed audience.
- config.go jwtAudiencesFromEnv() now ALWAYS unions BrandAudiences() into the
resolved allowlist (baked like the brand issuers). A legacy hanzo-only
GATEWAY_ALLOWED_AUDIENCES env override still accepts lux-cloud — the brand auds
don't depend on getting the deploy env perfectly right. Fail-secure: only ADDS
the known-good <brand>-cloud client_ids, never an arbitrary aud. unionStrings
dedupes so an env-supplied entry is never duplicated.
Tests: BrandAudiences (registry-derived, covers every brand), jwtAudiencesFromEnv
brand-union (baked default AND a hanzo-only env override both accept lux-cloud, no
duplicate). Paired with hanzoai/ai#64 (the per-brand signin code exchange).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Widen identityValidator to a trusted issuer SET (primary UNION BrandIssuers) so one
cloud binary validates hanzo AND lux/zoo/pars tokens off the one shared IAM JWKS.
Fail-secure: only known-good brand issuers added. NOT built (phantom go-sqlite3
v2.0.3 dep blocks) / NOT gated / NOT deployed — needs: per-brand EXCHANGE client
verification, build-dep fix, hanzo-login no-regression gate, red review.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
github.com/hanzoai/sqlite is the ONE Hanzo SQLite driver: it registers the
"sqlite" database/sql name under BOTH build tags (cgo → mattn+SQLCipher,
encrypted at rest; !cgo → pure-Go modernc, wrapped internally). Fourteen
cloud stores plus the pg→sqlite migration blank-imported modernc.org/sqlite
DIRECTLY, so a CGO_ENABLED=1 build registered "sqlite" twice — the fork's
mattn registration AND the direct modernc one — and panicked at init
("sql: Register called twice for driver sqlite"), taking down the whole
fused hanzo/cloud binary.
Swept every direct `_ "modernc.org/sqlite"` → `_ "github.com/hanzoai/sqlite"`
(sql.Open("sqlite", …) calls unchanged — same driver name). go.mod promotes
the fork to a direct require and demotes modernc to indirect (it survives
only as the fork's !cgo backend). Stale comments calling modernc the
"primary" driver corrected. Production already builds CGO_ENABLED=0 (one
modernc package, no collision); this makes the driver choice consistent and
unblocks a CGO_ENABLED=1 + libsqlcipher encrypted build.
NOTE: five UPSTREAM modules (base/core, ai/object, o11y sqlstore,
commerce/db, orm/db) still import modernc directly; a CGO_ENABLED=1 fused
build stays collision-prone until they adopt the fork too. Out of scope for
this repo; tracked as the cross-repo follow-up.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
First business app lane native on the Hanzo Framework DocType engine:
generic module-install path (/v1/framework/modules[/:module[/install]]) +
the CMS content model as fixtures (Page/Post/Article/Media/Navigation/Author,
module cms). Additive — no change to the proven /v1/framework/* isolation.
RED verdict: SHIP (0 crit/high/med).
The first Semgrep-class capability shipped natively in the cloud binary,
per hanzoai/security POSTURE.md's plan of record. One subsystem, the
established clients/ pattern (self-registering Mount, org-scoped store
under DataDir, audit + metering wired), zero external tools.
- engine.go — the reusable detection core: pure (path,content)→findings,
no I/O. Pattern rules (AWS/GCP/GitHub/Stripe/Slack/npm keys, private-key
blocks, JWTs) + a Shannon-entropy-gated generic-assignment rule so
`secret = "changeme"` is not flagged but a real high-entropy token is.
THE INVARIANT: a finding never carries the raw secret — only a masked
preview (4+4 ends, middle starred; short secrets fully starred) and the
SHA-256 fingerprint (dedupe + rotation tracking). Persisting plaintext
would make the findings DB the very thing we scan to prevent.
- store.go — per-tenant SQLite ({DataDir}/security.db), scans + findings
tables, org column is the isolation boundary on every query; SaveScan is
one transaction so a scan is never half-written. Mirrors clients/git.
- security.go — mounts /v1/security/{health,rules,scans,scans/:id,
findings,findings/:id}. submitScan runs the engine, persists redacted
findings, meters one unit, emits a tamper-evident audit record (the
tally, never the secrets). Registered cloud.RegisterWithShutdown(
"security", 136, …) + one blank-import line in subsystems.
Tests (14, no skips, no fakes): engine_test proves each rule fires, the
entropy gate, line mapping, dedupe, severity ordering, and that no field
ever echoes the raw secret; security_test proves the HTTP surface,
cross-tenant isolation (evil sees 0 of acme's scans/findings, 404 on id),
the no-principal 403, the severity filter, and that a clean scan persists
a real zero-findings record. go build ./... + go vet + -race all clean.
Co-authored-by: hanzo-dev <dev@hanzo.ai>
Framework: a generic, DRY module-install path so an app lane (CMS/ERP/Helpdesk)
declares its DocTypes as fixtures (framework.RegisterModule, sibling to
RegisterHook) and installs them per-org via the engine's own gate:
GET /v1/framework/modules list registered lanes
GET /v1/framework/modules/:module lane fixtures + which are installed in-org
POST /v1/framework/modules/:module/install ensure fixtures exist (managerOnly, idempotent)
'modules' reserved so the static routes are never shadowed by a document route.
CMS (clients/cms): the first lane — the content model as fixtures only, NO HTTP
surface of its own. Page/Post/Article (slug-named, status Draft/Published,
author Link), Media (Attach-backed DAM), Navigation (JSON menu), Author. A CMS
collection IS a framework DocType (module 'cms'); content IS documents;
publishing IS a status field. Registered at init; installed per-org.
Secure by default: install is managerOnly (owner seeded trust-on-first-use),
create-if-absent (never clobbers a customised DocType), stamps the module tag,
and every op stays per-org via principal.Tenant.
Tests: install/idempotency/unknown-404/tenant-isolation/forged-principal-403/
non-owner-403/module-tag; CMS fixture validity + content-model spec + a full
HTTP install->create Author->create Page(link)->publish->filter round-trip.
The router (zip over fasthttp) runs with Fiber's default UnescapePath:false, so
c.Param() returns path segments verbatim. A DocType/document/role name that is
legal per docTypeNameRe but contains a space ('Sales Invoice', 'System Manager')
arrives percent-encoded ('%20') and never matched its stored value: GET/PUT/
DELETE /v1/framework/:doctype/:name, submit/cancel, and revokeRole (:user/:role)
all 404'd, while create+list (no name in the path) worked — records could be made
yet be unreachable, and a granted System Manager could not be revoked.
Fix scoped to the ONE seam: pathParam() percent-decodes every framework path
param (getDocType/replaceDocType/deleteDocType, access(:doctype), docName,
revokeRole). NOT a global fiber.Config UnescapePath flip — that would change
segment splitting on the KMS secret-path, model-catalog, git and s3 wildcards
(c.Params("*")) that legitimately carry encoded slashes; the framework-local
decode is orthogonal and zero-blast-radius. Malformed escapes fall through to an
honest 404, never a panic.
Unblocks space-named DocTypes for the CMS/ERP/Help app lanes. Tests: red→green
round-trip (create -> GET/PUT/submit/cancel/DELETE by name) + space-named role
revoke; full framework suite green under -race.
Reconciles task #52 (eliminate Mongo) onto main's dedicated-per-org instance
model. The managed "document database" (docdb) is now a dedicated per-org Hanzo
Base instance — JSON document collections on per-tenant SQLite with native
realtime (SSE /v1/realtime), IAM-native — NOT a per-org FerretDB/Mongo instance.
- dedicated.go: docdb engine swapped ghcr.io/hanzoai/docdb:0.1.0 (mongodb://,:27017,
POSTGRES_*) -> ghcr.io/hanzoai/base (http://.../v1, :8090, /data), dsType "base".
New engine fields: dataMount (emits spec.volumeMounts so the data PVC actually
mounts — the operator does NOT auto-mount) and iamAuth (IAM-native: no per-
resource password generated/sealed/returned; admin Secret carries IAM_URL/
KMS_URL/IAM_CLIENT_* from the cloud binary's own IAM identity). baseInstanceEnv
helper. createDedicated honors iamAuth (no pw path). datastoreCR emits
volumeMounts. Runs on the operator's GENERIC Datastore controller — spec.type is
free-form, image/ports/volumeMounts drive the StatefulSet verbatim, NO operator
Rust change (verified). datastore engine stays ClickHouse (not Mongo).
- provisioning.go: header + sanitizeIdent comments de-Mongo'd.
- go.mod: mongo-driver demoted direct -> // indirect (zero Go imports of
go.mongodb.org remain; it survives only as a transitive requirement).
- test: TestDedicated_DocdbIsBase asserts the docdb CR is the base image on :8090
with /data mounted, connString http://.../v1 (no mongodb://), no credential
(IAM-native), IAM env in the admin Secret. PASS.
Audit (unchanged): zero customer docdb data, so the swap is clean.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Red found the self-service enablement write path (POST /v1/enablement/optin|optout)
+ the view keyed on raw c.Org() instead of principal.Tenant(c). On the bearer-less
direct-to-pod path SanitizeIdentity restores a client X-Org-Id with no validated
principal, so an off-gateway caller could opt an org it does not own into/out of a
beta (cross-tenant enablement write; bounded — betaOrgs membership of already-beta
items only, no money/data/global-state, not reachable through the gateway).
Fix (DRY — pricing already imports principal, the READ catalog gate already uses it):
- enablementOpt: resolve subject via principal.Tenant(c), 401 on !ok (was raw c.Org()).
- enablementView: resolve via the existing trustedOrg(c) (validated-principal gate).
- Corrected the docstrings (subject is the VALIDATED tenant, not 'never client-supplied').
Red's two PoC attack tests (enablement_attack_test.go) now GREEN; full enablement
suite unchanged + green + -race. Closes the one MEDIUM from Red's cockpit review.
Live-verified: /v1/billing/transactions returns {count, transactions:[...]}, not a bare array.
My client decoded a bare [] (test fake also returned bare — mock hid the bug), so the analytics
retention/churn/active/usage ledger read got ZERO rows despite real usage (maxpower: 268 txns).
Now decodes the wrapped shape (bare-array fallback for robustness); test fake mirrors the live shape.
User app secret env is no longer refused — it is sealed into cloud's embedded
KMS and wired to the pod via an operator-materialized k8s Secret, never
plaintext, never logged.
The path a secret takes (secrets.go):
1. SEAL — createApp / PUT .../env seal every secret:true value into deps.KMS
at a per-tenant/app coordinate (platform/<tenant-ns>/<app>/<KEY>);
the persisted env_json value is blanked. Fails CLOSED if KMS is
unavailable — a plaintext secret never lands in the DB as a fallback.
2. DECLARE — on deploy (applyLive, the ONE shared choke point) cloud writes a
canonical KMSSecret CR (secrets.lux.network/v1alpha1) into
tenant-<org> declaring a managed Secret <app>-env sourced from that
KMS scope. Best-effort: a missing CRD/RBAC degrades to an honest
'pending' status, never a failed deploy.
3. MOUNT — the Service CR renders each secret env as valueFrom.secretKeyRef →
that Secret (optional:true so the pod boots pre-sync); the hanzo
operator (which already supports secretKeyRef) mounts it into the
Deployment env. cloud is never in the plaintext path at runtime.
Also: PUT /v1/platform/projects/:p/apps/:a/env to set/rotate env post-create;
honest secretSync status (pending|syncing|ready|failed) from the KMSSecret CR
conditions on the app view; KMSSecret teardown on app/project delete.
Tests: seal blanks+seals+fails-closed; injective KMS refs (no cross-tenant
collision); canonical KMSSecret CR shape; apply/patch/delete; sync-status
mapping; and an end-to-end deploy asserting the Service CR carries secretKeyRef
(never plaintext) and the KMSSecret CR is authored. go build ./...=0, vet clean.
datastore (ClickHouse) + docdb (FerretDB) were honest-gated (unavailableKinds)
because a SHARED backend can't scope a per-tenant role. Replace the gate with a
DEDICATED-instance strategy: each create launches the org's OWN instance via an
operator Datastore CR + admin Secret in tenant-<org> (derived from the VALIDATED
org, never a request field). Isolation is BY INSTANCE — a cross-tenant grant is
impossible, there being one tenant on the instance — which un-gates both kinds.
- dedicated.go: engine table (image/ports/admin-env/DSN per kind), instanceName
(<prefix>-<orgHash10>-<name>, DNS-1123), the k8s orchestrator (ensure tenant
ns + RBAC wait, apply/observe/delete the Datastore CR + admin Secret + reap the
retained PVC) behind an interface a fake stands in for, createDedicated,
reconcileDedicated (provisioning->ready off the operator's status.phase), and
dropDedicated.
- Billing (first-class): a provision debit carrying the size dimension
(Model=<kind>:<size>) lands on the CALLER's org via the ONE commerce meter, and
a recurring GB-day footprint sweep charges every running instance's own org —
the reserved hook, now unblocked by the instance's declared size. Drop removes
the row, stopping the meter, and reaps the PVC so no storage leaks.
- unavailableKinds now empty (mechanism kept); shared datastore/docdb
provisioners deleted (one way only); the 5 shared kinds untouched.
- Datastore CR (not DocDB) is used because only its controller writes
status.phase, the readiness signal; type forced per engine.
Tests: hermetic dedicated suite (fake orch + mock commerce) proves CR/Secret
shape, two-org isolation, ready reconcile, drop+PVC reap, and per-org billing
attribution; a build-tagged livecluster test proves the whole path against the
real operator. go build ./cmd/cloud + go test ./clients/provisioning/... green.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Completes the console subsystem's standalone-route port for the True 1-binary
FE. The three remaining console2 Next server routes that do REAL server work
(not vanishing BFF proxies) now terminate natively in the unified binary, so
console2 can drop them from its static export:
POST /v1/console/waitlist waitlist.go — session-gated join to the Base
waitlist plugin; the recorded email is BOUND to
the gateway-verified X-User-Email (a signed-in
user can't enroll a third party), honest 501
when WAITLIST_URL is unset.
GET /v1/console/embed-status embed.go — server-authoritative entitlement
(owning brand org / global admin only) + a
time-boxed reachability probe. SSRF-free: the
target is <app>.<brand-domain> for the FIXED
deployment brand (deps.Brand) — no client host
in the target at all.
POST /v1/console/topup/wallet topup.go — verify an HUSD transfer on-chain
(plain eth JSON-RPC, no EVM dep) and credit the
VALIDATED caller's own org for the ON-CHAIN
amount via the S2S commerce billing API. IDOR-
safe (ignores any client userId); honest 501
greenfield gate until HUSD is deployed.
All three resolve the caller from the VALIDATED principal only (same trust
boundary as keys/onboard); a forged X-User-Id/X-Org-Id is refused. docs is a
pure host->URL redirect with no server work, so it stays client-side in console2
(no handler here). Registered in the ONE routes() place; full unit coverage
(fake IAM/waitlist/RPC/commerce), go build ./... clean, binary boot-proven to
serve the console SPA at / with every /v1/console/* route resolving (403/501/503,
never 404).
v1.786.33 set zip.Config.ReadBufferSize=32768 (GATEWAY_READ_BUFFER_SIZE) but
zip v1.2.0's HTTP transport built a bare fasthttp.Server and dropped it — the
edge still 431'd at 4 KiB. zip v1.2.1 propagates the App Config onto the
transport's fasthttp.Server, so the 32 KiB header ceiling now takes effect on
cloud:8000 (the api.hanzo.ai/v1/* backend).
The console E2E saw /v1/crm/summary miss a just-created record. Root cause was a
stale/eventually-consistent read; the current handler already counts LIVE
(s.Counts -> SELECT COUNT(*) per table on the same store the writes hit), so a
create/delete is reflected with ZERO lag — verified live (create company ->
summary companies +1 immediately).
Add TestSummaryReflectsCreateImmediately: create -> Counts shows +1, delete ->
Counts shows -1, all in one synchronous flow. This guards against any regression
to a materialized/async rollup. No production code change needed — the fix is the
regression lock.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
Gap 2 (analytics 502 on ClickHouse i/o timeout): /v1/analytics/{overview,timeseries,top}
returned a raw 502 when a DatastoreQuery hit a connectivity failure even though
requireDatastore()'s not-connected path already 503s. warehouseErr() now maps a
transport/connectivity error -> 503 'warehouse unavailable' (retryable); a REACHABLE
warehouse that rejected the query (bad SQL/protocol) stays 502. Table-driven test.
Gap 3 (docdb/datastore provisioning) — SAFETY REWORK (replaces the earlier
readWriteAnyDatabase change, which was cluster-wide = a cross-tenant hole):
- datastore + docdb are honest-GATED (unavailableKinds -> 503 'not yet available',
refused BEFORE billing or any backend write) because their backends cannot mint a
per-tenant-SAFE credential:
* datastore (ClickHouse): no grant-capable per-tenant admin (GRANT ALL -> Code 497);
unblocking is backend-side (StatefulSet grant-capable admin).
* docdb (FerretDB/DocumentDB): engine implements ONLY cluster-wide roles
(clusterAdmin -> Postgres SUPERUSER, readWriteAnyDatabase) — no per-db role.
- docdbProvisioner.Create keeps requesting the CORRECT per-db 'readWrite' role (the
tenant-safe target); when FerretDB supports it, drop the gate and it works as-is.
- Tests: gated kinds -> 503 with the provisioner never run; the 5 guaranteed kinds
(sql/vector/kv/search/s3) are asserted NOT gated.
Bar met: 5/7 data kinds fully work; datastore+docdb show an honest 'coming soon',
never a security hole.
Co-authored-by: Hanzo Dev <dev@hanzo.ai>
The public HTTP edge (zip/fiber) uses fasthttp's default 4 KiB per-conn
read buffer, which caps total request-header size and returns 431 (Request
Header Fields Too Large) above it. Once an admin-guard Domain=.hanzo.ai SSO
cookie is set on every subdomain, a browser's request headers cross ~4 KiB
and every request to api.hanzo.ai/v1/* (gateway -> cloud passthrough) 431s.
Raise the edge ceiling to a sane 32 KiB (nginx large_client_header_buffers
parity) via zip.Config.ReadBufferSize, env GATEWAY_READ_BUFFER_SIZE (shared
with the gateway edge so both trust boundaries agree on ONE value; tunable
down if the per-conn memory budget demands). Internal zip services keep the
4 KiB framework default — only the browser-facing edge opts up.
Repro (pre-fix): POST cloud:8000/v1/agents with a 9 KiB Cookie -> 431
Server: fasthttp. Post-fix: same request -> 403 (auth), no 431.
create and list return an agent's public id (agent_<hex>), but get/run/update/
delete/runs resolved the URL path segment ONLY against the name column — so a
client that used the id create returned got 404 "agent not found". A created
agent was listed but neither gettable nor runnable by the identifier the API
handed back.
One-way fix: Store.Resolve(org, ref) matches either the public id or the
org-unique name (id wins on the astronomically-unlikely in-org collision),
org-scoped fail-closed so a cross-tenant ref is 404, never a leak. Every
path-addressed handler (get/update/delete/run/runs) resolves through it and keys
every downstream store op on the resolved a.Name. Route param :name -> :ref to
say what it accepts. The run path's validated-principal gate and single
product=agent debit are unchanged.
Tests (go test -race): Resolve by id and by name return the SAME agent; full
create -> get-by-returned-id -> run-by-returned-id all 200 the same agent with
real output; cross-org ref denied 404; a run addressed by the returned id meters
exactly once (product=agent).
The 1-binary console (HIP-0106) shipped as a 3-file STUB: nothing ran
console2's static export before `go build`/the image build, so `//go:embed
all:webui/dist` baked only the fallback shell.
- `make webui` (new): runs hanzoai/console2 `npm run build:embed` and overlays
the static export into webui/dist so a plain `go build` embeds the full
@hanzo/gui console. `make build-standalone` = webui → build. CONSOLE2_DIR
points at a console2 checkout (default ../console2).
- Dockerfile console stage + webui.go + the stub shell: docs corrected — the
pipeline is `build:embed` (a static export), not `next build`; the export now
prerenders clean (console2 build-embed.mjs neutralizes the root layout's
request-time headers() read), so the image embeds the real console instead of
silently degrading to the shell. Bumped the export heap to 8192 for headroom.
Verified: build:embed → webui/dist (index.html 368 KB, /_next/static assets) →
CGO_ENABLED=0 go build ./cmd/cloud → the running binary serves the real console
at / (200, references /_next/, not the stub), the SPA shell for deep links
(/orgs), fingerprinted assets immutable-cached, and the /v1 API on the SAME
origin ({"service":"base","status":"ok"}); unmatched /v1/* is a real 404, not
HTML. webui_test.go (7 tests) green against the real bundle.
Red measured 3-6 System Managers seeded when concurrent role-less members first
administered a fresh org: managerOnly did a check-then-insert (OrgHasRoles then
AssignRole) with a TOCTOU window. Fix: store.SeedOwnerIfUnowned is a SINGLE
conditional INSERT ... SELECT ... WHERE NOT EXISTS(SELECT 1 FROM fw_roles WHERE
org=?), so the unowned-check and the insert are one atomic statement — exactly
one concurrent first-caller's row lands. RowsAffected==1 => this caller is the
seeded owner; ==0 => re-resolve (a concurrent grant may have made them a
manager) else 403. No UNIQUE index (multiple SMs are legit later via AssignRole;
only the AUTO first-seed must be singular). Removed the now-dead OrgHasRoles.
Test: TestAtomicOwnerSeed — 8 concurrent role-less first-callers → exactly 1
seeded winner + exactly 1 System Manager row. 22 tests total, race-clean.
searchGuard treated the searxng X-API-Key as OPTIONAL — a MISSING key passed —
so GET /v1/websearch/search was an open proxy to the Hanzo-operated metasearch
instance (unauthenticated request-forgery + cost surface). Its scrape sibling
(scrapeHandler) already fails closed; this brings search to parity:
- key unset → 503 (surface not configured, never open-to-all)
- X-API-Key missing → 401 (constant-time compare of "" vs want fails)
- X-API-Key mismatch→ 401
Safe for the real caller: the LibreChat searxng client sends the configured
searxngApiKey (universe chat configmap wires searxngApiKey=${WEBSEARCH_API_KEY})
as X-API-Key, so only anonymous callers are turned away.
Tests: TestSearchMissingKeyRejected (was ...Allowed) → 401; new
TestSearchUnsetKeyFailsClosed → 503; TestSearchProxyRewritesToSearchPath and
TestMountRoutesThroughRouter now present the key. go build/vet/test green.
LOW-1 — Single submit-immutability: updateDocument/createDocument for a Single
now route through writeSingle, which enforces the SAME draft-only guard as the
non-Single path (a submitted/cancelled Single → 409, not a silent mutation) and
preserves a redacted Password across an unchanged update.
LOW-2 — secure-by-default permissions (no open-to-all footgun):
- permission.can() is now DEFAULT-CLOSED: removed the 'empty perms => open to
every org member' branch. A permless doctype is manager-only; a role-less
member is denied.
- DocType.normalize() seeds a System Manager perm at define time, so a stored
doctype is never silently permless (explicit in UI + audit).
- Owner seeding moved from resolveAccess (any member is SM until a role exists)
to managerOnly as trust-on-first-use: the FIRST validated principal to
administer an org with no roles becomes its persisted System Manager (the
owner) — exactly one member, deterministically, never cross-tenant.
Tests: +2 (TestSingleSubmitImmutability, TestPermlessDefaultClosed); 21 total
race-clean. go build ./... CGO=1 & =0 green, vet + gofmt clean. Fixed binary
boot-verified (framework health 200, gate 403, forge 403).
The prior fix re-fetched luxfi/age with checksum-checking off, which recorded the
DIRECT-vcs hash. luxfi/age@v1.5.0 was force-re-tagged, so the direct tree hash
differs from the immutable proxy zip hash — the container build (GOPROXY=proxy +
GOSUMDB=sum.golang.org, GOPRIVATE dropped for luxfi/* on purpose) verifies against
the sumdb and hit "checksum mismatch / SECURITY ERROR" on `go mod download`.
Replaced the h1: zip hash with the canonical value from
sum.golang.org/lookup/github.com/luxfi/age@v1.5.0 (the /go.mod hash already
matched). Now go.sum == what the proxy+sumdb serve → container verification passes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The console-embed stage's contract is "a missing static target is a degrade, not
an error" — but the guard only handled the target being ABSENT. When console2
exposes build:embed AND it CRASHES (currently: /signin Server-Components
prerender error kills `next build`), the `&&` chain failed the whole cloud image,
so a frontend prerender bug took down the entire Go backend build (release runs
for projectsvc S3 fix + kms refactors all failed here, not on Go).
Complete the stated contract: wrap build:embed so a build FAILURE also degrades to
the committed fallback shell (/out stays empty → Go embeds webui/dist/index.html).
The standalone console2 Deployment is the primary console; this embed is a
same-origin convenience and must never gate the backend image.
(console2 /signin static-export prerender crash tracked separately for the
console track — this makes cloud CI robust to it either way.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two dep-rot issues blocking the cloud (Go backend) build:
1. A transitive dep requires the non-existent mattn/go-sqlite3 v2.0.3+incompatible.
The unversioned replace didn't stop Go reading v2.0.3's go.mod during graph
load. Fixed with a VERSIONED replace (v2.0.3+incompatible => v1.14.16, the last
real go-sqlite3, drop-in package sqlite3). cloud's primary sqlite is
modernc.org/sqlite (pure-Go); hanzoai/sqlite (encrypted, package `sqlite`) is a
separate driver, adopting it is a real migration not this phantom fix.
2. luxfi/age@v1.5.0 go.sum checksum mismatch → removed stale lines + re-fetched.
go build ./internal/org/ (the sqlite consumer) now clean; module graph resolves.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The Hanzo Framework: Frappe's DocType/metadata core rebuilt native in Go on
Base/SQLite, mounted at /v1/framework/* (subsystem order 129). ONE engine +
ONE generic UI renders every business app — CMS content-types, ERPNext
DocTypes, Helpdesk all become just DocTypes on this engine. No Frappe/Python
runtime dependency; the engine is pure Go.
- DocType registry: define/list/get/replace/delete metadata per-org
- Generic metadata-driven document CRUD with ?filters=/fields=/order_by=/limit=
- Fieldtypes: Data/Int/Float/Currency/Check/Date/Datetime/Text/SmallText/
LongText/Select/Link/Table/Attach/JSON/Password (all validated)
- Naming: hash / field: / prompt / series patterns (INV-.YYYY.-.#####)
- Link relations (in-org ref check + fetch_from), Table child rows
- docstatus 0/1/2 with submit/cancel for submittable doctypes
- Per-org permissions (DocType perms by role) + per-org role store
- Go lifecycle hook interface (before_insert/before_save/after_save/
on_submit/on_cancel/on_trash) — gpython/goja runner is a later add to the
SAME interface
- Password fields: argon2id hash on write, redact on read (fail-secure)
Security: org derived ONCE via clients/principal.Tenant (validated principal
only; forged X-Org-Id refused 403). Every table + query is org-scoped. 19
tests (race-clean) prove cross-org isolation, forged-principal refusal,
permission enforcement, field-type validation, and the docstatus lifecycle.
Boot-verified locally (health 200, doctypes 403, forge refused).
The embedded luxfi/kms core (cloud/types-only leaf, built by build.go before the
app exists to break the import cycle) is now just 'kms'; the Fiber /v1/kms/* mount
subsystem (imports cloud) is 'kmssvc' (its existing internal name). One clean name
each, no unnecessary compound. build+vet+tests green.
publicReadPolicy used Principal {"AWS":["*"]} + array Resource, which
SeaweedFS's S3 policy engine rejects with 'Policy has invalid resource' —
aborting ensureBucket (SetBucketPolicy) BEFORE any files upload, so every
projectsvc deploy failed ('object storage'/'invalid resource') and no site
was ever served. Use scalar Principal "*" + scalar Resource, which SeaweedFS
accepts and is equally valid on AWS S3 / MinIO. Verified: mc anonymous
set-json with this exact shape succeeds against the s3.hanzo.ai SeaweedFS
gateway; a site uploaded to the now-public hanzo-sites bucket serves 200 at
https://s3.hanzo.ai/hanzo-sites/<org>/<slug>/index.html.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Found by testing the ACTUALLY-deployed hanzoai/crawl:0.0.1 (= Crawl4AI
0.8.6) against clients/websearch's crawl adapter: 0.8.x/0.9.x return the
/crawl result's `markdown` as an OBJECT
{raw_markdown, fit_markdown, markdown_with_citations, ...}, and signal the
batch with a boolean `success` (no `status`). crawlResult.Markdown was
typed `string`, so json.Decode errored on the object form → crawl()
returned an error → EVERY scrape returned {success:false} with empty
content. hanzo.chat Web Search's scrape half was therefore dead even once
crawl is running.
Fix: markdownField.UnmarshalJSON accepts either a bare string OR the
object (preferring the cleaned fit_markdown, raw_markdown fallback); the
crawlResponse envelope now also accepts boolean `success` alongside the
legacy `status`. Neither envelope field is required — Results[0].Success
is authoritative.
Tests (real 0.8.6 response shape):
TestScrapeHandlesCrawl4AIObjectMarkdown — object markdown + bool success
→ success:true, returns fit_markdown (was: {success:false}).
TestMarkdownFieldAcceptsBareString — bare-string form still works.
All 12 clients/websearch tests pass; go build + go vet clean.
Contract verified live: crawl4ai 0.8.6 POST /crawl {urls:[...]} returns
synchronously (no task_id polling) with url/markdown/success/metadata —
matches the adapter otherwise.
deploymentLogs returned only the recorded timeline + a Job reference, ending
with a '(live BuildKit Job logs stream in phase 2)' placeholder. This closes
that phase-2 gap: it streams the ACTUAL pod logs from the cluster — the BuildKit
Job's pod while a git build runs, and the running app's pod once deployed — so
the console's per-deployment Logs pane shows real output, operator-consistent.
- logs.go: buildLogs (job-name=<jobName> pod in the build ns), appLogs
(app.kubernetes.io/instance=<slug> pod in tenant-<org>), one podLogsBySelector
path (newest pod, tail-bounded 400 lines, byte-capped 256 KiB keeping the tail,
8s time-boxed). Every read is org-scoped and time-boxed.
- k8s.go: add a typed kubernetes.Interface clientset (from the SAME rest.Config)
held ONLY for the Pods().GetLogs subresource the dynamic client cannot express;
nil-safe — a construct failure leaves logs degrading to the timeline and never
disables the CR control plane (all on dyn).
- deploy.go: deploymentLogs now appends real build + app logs and stamps a
(build|app|none) so the console can label the pane. HONEST DEGRADE:
an unreachable cluster / absent pod yields the recorded timeline + a stated
'not available' note — source stays reflecting real streamed content, never a
fabrication.
- 8 tests over a fake typed clientset: newest-pod selection, tenant-namespace
scoping (acme never reads victim's pod), no-pod/no-clientset honest degrade,
and the handler surfacing live build logs (source=build) vs degrading
(source=none) — asserting the phase-2 placeholder is gone.
The console can be a static export only once its two NON-proxy Next server
routes (app/keys, app/onboard) — which mint/revoke the user's hk- Cloud API
key and create the user's org as the confidential hanzo-console IAM client —
have a native home. Port them to a clients/console subsystem mounted at
/v1/console/* in the one binary (task #41, True 1-binary FE): the embedded SPA
calls /v1/console/* on its own origin, and the last stateful Node handlers go.
- clients/console/iam.go: confidential-client (client_secret_basic) IAM caller
for mint/revoke/get the hk- key + create/read/update an org. Honest 501 when
IAM_MINT_CLIENT_ID/SECRET are unset (mirrors identity.ts mintConfigured()).
- clients/console/console.go: /v1/console/{keys(GET/POST/DELETE),onboard(POST),
health}. Every route requires a VALIDATED principal; the IAM id is DERIVED as
<owner>/<name> from the gateway-minted X-User-Id/X-Org-Id, never a request
value — a caller can only ever act on their OWN key/org (red-bar structural).
- clients/console/onboarding.go: faithful Go port of console2 onboarding.ts —
pure slug + reserved-name policy (admin/built-in/app + hanzo/lux/zoo/pars).
- Registered as consolesvc (order 122) so /v1/consolesvc/health does not shadow
the real fail-closed /v1/console/health probe.
- 16 tests: unauth 403 (forged X-Org-Id refused, IAM never touched), mint/get/
revoke scoped to the derived id + show-once + no secret leak on GET, 501/502
honesty, onboard first-run(create+move)/additional(create-only)/reserved 400/
taken 409/personal auto-suffix, + pure-policy unit tests.
Close the two PaaS domain gaps so a customer can put their app on their own
domain, operator-native, from the console.
- Seed a canonical default host <slug>.<org>.<sitesHost> on app create, so
every app has a working HTTPS URL the moment it deploys (operator issues the
cert). Never removable.
- New /v1/platform/.../domains surface (list/add/verify/remove). An org-subtree
host is active immediately; a BYO custom host (yourco.com) is claimed PENDING
and returns the exact DNS records to publish (TXT ownership token at
_hanzo-challenge.<host> + CNAME to the app host).
- Verify resolves DNS: a matching TXT token proves control (DNS-01 model), then
the host is rendered into the app's operator Service CR ingress via applyIngress
(cert-manager TLS comes for free). Honest still-pending on not-yet, never fake.
- platform_domains table: host PRIMARY KEY = global uniqueness (one org per host,
like site_hosts); pending→verified lifecycle. Cascade-deleted with app/project.
- validateOrgDomains extended: a non-subtree host renders ONLY when this org owns
a VERIFIED claim; unverified/foreign/apex hosts still refused (RED hardening kept).
- ingressSpec extracted (one TLS shape shared by serviceCR + applyIngress);
observeDomains surfaces operator status.endpoints/phase for honest live state.
- Tests: verified-custom accept + pending/foreign refuse; full add→verify→remove
HTTP flow with fake DNS; global uniqueness (two orgs/two apps); apex refusal;
default-host seeding; CR ingress render. go build + go test green.
Found by testing the ACTUALLY-deployed hanzoai/crawl:0.0.1 (= Crawl4AI
0.8.6) against clients/websearch's crawl adapter: 0.8.x/0.9.x return the
/crawl result's `markdown` as an OBJECT
{raw_markdown, fit_markdown, markdown_with_citations, ...}, and signal the
batch with a boolean `success` (no `status`). crawlResult.Markdown was
typed `string`, so json.Decode errored on the object form → crawl()
returned an error → EVERY scrape returned {success:false} with empty
content. hanzo.chat Web Search's scrape half was therefore dead even once
crawl is running.
Fix: markdownField.UnmarshalJSON accepts either a bare string OR the
object (preferring the cleaned fit_markdown, raw_markdown fallback); the
crawlResponse envelope now also accepts boolean `success` alongside the
legacy `status`. Neither envelope field is required — Results[0].Success
is authoritative.
Tests (real 0.8.6 response shape):
TestScrapeHandlesCrawl4AIObjectMarkdown — object markdown + bool success
→ success:true, returns fit_markdown (was: {success:false}).
TestMarkdownFieldAcceptsBareString — bare-string form still works.
All 12 clients/websearch tests pass; go build + go vet clean.
Contract verified live: crawl4ai 0.8.6 POST /crawl {urls:[...]} returns
synchronously (no task_id polling) with url/markdown/success/metadata —
matches the adapter otherwise.
Front the Lux chain-data plane over HTTP so the console's Indexer and
Oracles pages read REAL chain state from api.hanzo.ai/v1/* instead of
rendering "not connected":
- GET /v1/indexers -> luxfi/indexer explorer REST (/health + latest
block): per-network chain/network/height/health. lag honestly omitted
(the indexer REST exposes indexed height, not the chain HEAD).
- GET /v1/oracles -> luxfi/graph GraphQL priceFeeds (O-Chain PriceFeed
registry): real on-chain price feeds; honest-empty when none.
Principal-gated (403 without a validated IAM principal); brand-scoped
(each brand's cloud is wired to its own indexer/graph, a ledger is public
within a brand). Honest 502 on unreachable upstream, never a fabricated
row. Mirrors clients/visor + clients/zt structure; interface-seam tests
against a fake upstream. Registered order 135.
Env-gated on OTEL_EXPORTER_OTLP_ENDPOINT; non-fatal; clean no-op when unset (safe to ship before the collector is live). Installs the global tracer provider with a service.name resource so the console Monitoring tab filters this product. Mirrors ai/object/telemetry.go. Traces-only; metrics/logs are a tracked follow-up.
RED review fixes on the sites router:
1) [HIGH] ONE reserved-subdomain source (clients/sites/reserved.go:
baseReserved baked-in + operator SetReservedExtra, never subtractable),
consulted at THREE points that can no longer drift: serve (siteSlug),
project-create (createProject -> 400), and host-bind (Store.BindHost ->
errReservedHost). site_hosts can now NEVER physically hold a reserved host,
so a reserved subdomain never resolves even if the ingress regex drifts —
the serve gate is a backstop, not the sole guard. Widened the set to app/
auth/payment/brand labels (console, sites, internal, gateway, login, secure,
account, signin, auth, pay, wallet, admin, brand terms, ...).
2) [MED DoS] Serve now STREAMS objects (Fiber SendStream, Content-Length from
info.Size, fasthttp closes the reader) instead of io.ReadAll-buffering up to
64 MiB per request on the unauthenticated edge — removes the OOM vector.
Same for the 404.html path.
3) [LOW] Non-GET/HEAD on a site host → 405 + Allow: GET, HEAD.
Tests: IsReserved, reserved-host-never-serves backstop, 405, BindHost-rejects-
reserved (even with a forced project row), create-rejects-reserved-slug via the
real handler. All green; no regressions.
Close the two LOW follow-ups on the cold-start tenant-RBAC fix, plus confirm
the fresh-org fail-closed status. All on top of v1.786.23 (already SHIP).
L2 — git reconciler no longer treats a transient tenant-RBAC delay as TERMINAL
and no longer head-of-line-blocks other orgs. reconcileBuild now does ONE
non-blocking readiness probe (ensureTenantReady: create-namespace-if-absent +
single SelfSubjectAccessReview) instead of the synchronous image path's ~45s
in-line waitForTenantRBAC. If the operator's RoleBinding has not landed, the
deployment stays 'building' and re-drives on the next 10s tick — never a
permanent fail (there is no client to retry a git build) and never a 45s stall
of the shared sequential reconciler. Only the elapsed build deadline fails it
honestly. errTenantProvisioning from applyLive is also caught as transient
(defense in depth). Namespace-create is decomplected into one shared
ensureNamespaceExists; ensureNamespace (sync, blocking) and ensureTenantReady
(async, probing) compose it.
L1 — image deploy path gains a per-org in-flight-deploy cap (inflightGate,
maxConcurrentDeploys, default 8 via CLOUD_PLATFORM_MAX_CONCURRENT_DEPLOYS),
mirroring the git build cap. deployImage acquires a slot before applyLive's
~45s RBAC wait and releases on any return; over-cap is a retryable 429 refused
BEFORE recording an attempt. Bounds request-goroutine pile-up on a wedged
operator; per-org (one org's saturation never throttles another); fail-closed.
I2 — confirmed the truly-fresh-org path fails closed on 503 (RBAC pending),
NOT a raw 502: the namespace IS created (the trigger for the operator's
RoleBinding), then the bounded RBAC wait yields errTenantProvisioning -> 503.
Tests (all -race green): reconciler stays 'building' then goes live on a later
tick once RBAC lands (not 'failed'); over-cap image deploy -> 429 with per-org
isolation + slot-release re-admit; fresh-org deploy -> 503 with namespace
created + no Service CR + honest 'error' deployment recorded.
The /v1/admin money panels (finance/orgs/overview) read $0 for every org despite
real balances (lux $10,000, maxpower $20,498) because the commerce client used
the wrong org selector on BOTH axes:
- commerce.go get(): sent X-IAM-Org-Id, which commerce does NOT read. Commerce
EdgeAuth resolves the per-org billing namespace from the TRUSTED X-Org-Id header
(trusted only with the COMMERCE_SERVICE_TOKEN bearer). X-IAM-Org-Id silently
fell back to the default (COMMERCE_SERVICE_ORG) namespace.
- admin.go orgSubject(): keyed the wallet subject as "org/org"; commerce keys the
per-org wallet under the BARE org slug (user=<org>) within the X-Org-Id namespace
(the 2026-07 commerce durability rework, commerce >=1.46.8).
Either alone zeroed the reconciliation; both were present. The prior comments
encoded the wrong model ("commerce resolves from COMMERCE_SERVICE_ORG, header
advisory") — corrected to the verified contract.
Verified LIVE against commerce /v1/billing/{balance,usage-rollup}:
user=lux + X-Org-Id: lux -> $10,000.00 (1,000,000c)
user=maxpower + X-Org-Id: maxpower -> $20,498.13 (2,049,813c)
user=lux/lux OR X-IAM-Org-Id -> $0 (the bug)
The fleet-wide /v1/costs COGS god-view is org-independent and correctly sends no
org (unchanged).
Regression guard: TestCommerce_ReconcilesWithXOrgIdBareSlug — a contract-accurate
fake commerce that returns money ONLY for X-Org-Id + bare-slug user; proven
red->green (fails on org/org, passes on the fix). Full admin suite green.
Co-authored-by: blue <blue@hanzo.ai>
Red review of the live agent-session control plane.
FIX (MEDIUM, systems lens): sessionsStream retained root := c.Query("root")
verbatim. c.Query is a zero-copy view into the fasthttp request buffer, and the
SendStreamWriter loop OUTLIVES the handler (runs after the Ctx is recycled), so
the long-lived root filter raced a reused buffer — within-org stream-filter
corruption / UB. tenant() already clones org for this exact reason; clone root
the same way. Not cross-tenant (org is cloned + bus-filtered); fixes the race
and honors the file's own 'never touch Ctx after return' invariant.
TEST (vector #4): add TestSessionEventSeqConcurrent — 64 parallel AppendEvents
to one session must yield seqs exactly {1..N}, no gaps (no lost write) no dupes
(no raced MAX+1). Proves the single-conn + UNIQUE(session_id,seq) guarantee
under -race instead of only asserting it.
The canonical cloud registry every surface hangs off: live agent SESSIONS +
the subagent tree, streamed over ZAP, remote-controllable. This is the
view/control/stream layer; durable execution rides hanzoai/tasks, not a
bespoke scheduler.
Model + store (agents.db, same tenancy pattern as agents/runs):
- Session{id,agent,org,actor,status,parentSessionId,rootSessionId,title,
startedAt,endedAt,taskWorkflowId,taskRunId,events[]}. Subagent tree =
sessions linked by parentSessionId; the outer agent is the root, each
spawned subagent a child, all sharing rootSessionId. Parent must exist
in the SAME org (TOCTOU-checked in the write path) so a tree can never
cross tenants. Per-session monotonic event Seq.
REST (org-scoped via principal.Tenant, fail-closed):
- POST /v1/agents/sessions register (opt parentSessionId)
- GET /v1/agents/sessions list (filter root/parent/status)
- GET /v1/agents/sessions/:id detail + children + recent events
- GET /v1/agents/sessions/:id/tree full subagent-flow graph (1 query)
- PATCH /v1/agents/sessions/:id status/title (terminal is monotonic)
- POST /v1/agents/sessions/:id/events append message/tool-call/spawn/log
- POST /v1/agents/sessions/:id/{pause,resume,stop,message} control
Routes register BEFORE /v1/agents/:name (Fiber matches in registration
order); /stream precedes /:id for the same reason.
ZAP live stream:
- GET /v1/agents/sessions/stream (SSE) rides the ZAP machine transport
natively (zip SendStreamWriter streams through ListenZAP — proven by
zip stream_test). In-process bus is the single fan-out seam a direct
ZAP push subscription attaches to. Org-filtered, non-blocking, laggard-
drop; GET endpoints are the source of truth.
Durable execution = hanzoai/tasks (architecture alignment):
- Root session -> a tasks workflow; subagent -> child workflow (same
rootSessionId). TaskController seam mirrors the tasks SDK Client
(Signal/Cancel); control forwards to it when a session is task-backed,
else records the command as a durable control event for stream-
consuming surfaces. Default is the disabled (record-only) controller;
the live client.Dial(TASKS_URL) plug-in point is marked in Mount.
Run integration (#5): the ONE runAgent path (HTTP + scheduler) opens a
root session per run (best-effort, never fails the run), so every run is
visible in the same registry.
Tests (real): store tree-linking + cross-tenant/dangling parent deny +
event seq/counts; HTTP tree assembly; cross-tenant read/tree/control/
append/parent deny; control authz (no validated principal -> 403) +
tasks forward (signal/cancel, forward-failure 502, record-only fallback);
event append/seq + status monotonicity; run-opens-session; bus fan-out/
org-filter/overrun/close. go build+vet+test clean; -race clean.
Add clients/sites: a HOST-routed public site server that turns
<slug>.hanzo.app into the static site a project deployed to OUR S3
(<org>/<slug>/ in CLOUD_PROJECTS_BUCKET). Installed as the FIRST middleware
in the compose root, ahead of identity/billing, so a published site is a
public artifact served straight from S3 — never a tenant API call.
Tenant isolation (RED-focus): the org + S3 prefix come ONLY from the store
lookup keyed by the validated subdomain slug, never from the request path or
a client header. Object keys are rooted-clean (path.Clean under '/') so no
../ or encoded traversal can escape the <org>/<slug>/ prefix into another
project or org. A globally-unique site_hosts binding table makes a bare
subdomain resolve deterministically to exactly one tenant (project slugs are
only org-unique); binding is first-come and cannot be hijacked.
Cache: one canonical policy (sites.CacheControlFor) applied both when writing
objects at deploy and when serving them — HTML public,max-age=60,s-maxage=86400;
content-hashed assets immutable 1y; middle TTL otherwise; per-project
cacheControl override on the document TTL. Cloudflare purge-by-cache-tag
(site-<org>-<slug>) on redeploy AND delete; creds from KMS/env
(CF_API_TOKEN/CF_ZONE_ID), honest no-op when unset. Cache state (TTL +
lastPurgeAt) exposed on the project API.
Tests: traversal/cross-tenant isolation proof, host-routing + reserved-host
exclusions, first-come/no-hijack subdomain binding, CF purge client.
Add clients/zt — a thin, org-scoped facade over the Hanzo Zero Trust
controller's OpenZiti Edge Management API (/edge/management/v1), backing
the console's Networks, Service Mesh and Edge pages (which render
"not connected" today).
Surface (all org-scoped by the validated principal):
GET /v1/networks[/:id] the org's ZT overlay, projected from its edge-routers
GET /v1/mesh/services ZT edge services
GET /v1/edge/nodes ZT edge-routers + real online/disabled/offline status
- client.go: one HTTP path — Ziti password-auth (KMS-injected
ZT_CLIENT_ID/ZT_CLIENT_SECRET, zt-session header), cached session with
re-auth-and-retry-once on 401, generic {data,meta} pager, honest error
mapping, TLS trust via ZT_CA_PEM. Fails closed 503 when unconfigured.
- types.go: ZT wire structs + console view structs + PURE mapping.
Tenant isolation is the "org-<org>" role attribute (the ONE tenancy
convention ZT expresses natively); list/get filter to the caller's org.
- zt.go: routes/handlers, registered as subsystem "zt" (order 134).
- http_test.go: fake controller (interface seam) — asserts 200, tenant
isolation, shape, health mapping, 401 re-auth retry, fail-closed 503.
Honest-empty over fabrication throughout: no org tag -> invisible; no
routers -> no network; no metrics -> omitted (UI renders em dash).
New global-admin read GET /v1/admin/compute powering the console Bots + Machines
operator boards. Aggregates hanzo.compute_usage(org, app, project, kind, event,
machine_id, size, price_cents, ts) grouped by (org, app, project, kind) over the
shared datastore client (aiobject.DatastoreQuery — the clients/analytics transport,
no second conn). `kind` is an OPEN LowCardinality spectrum (bot|machine|cluster|
nodepool|container|function|…) matched as a PLAIN STRING — ?kind= narrows to any
kind (Bots=bot, Machines=machine; future Clusters/Functions reuse this endpoint),
?org filters, ?range=24h|7d|30d bounds. Two-level roll-up: inner argMax(event,ts)
per machine -> outer counts machines, active (latest non-terminal), sum(price_cents),
max(ts). Honest-empty when the warehouse/table isn't wired yet (visor/commerce
emitter pending) — never a fabricated fleet. Global-admin only (s.guard); stays v1.x.x.
New clients/visor subsystem fronts Visor (the cloud OS at visor.hanzo.svc) and
serves the console's Machines/GPUs/Clusters pages as clean, tenant-scoped REST off
the unified cloud binary — replacing the god-mode /paas admin proxy that 501s.
Routes (every route org-scoped by the validated principal → Visor ?owner):
GET/POST /v1/machines, GET/DELETE /v1/machines/:id -> get-machines / machines/launch / delete-machine
GET /v1/gpus (+ /v1/gpus/alerts) -> per-accelerator inventory derived from GPU machines
GET /v1/clusters, node-pool create/scale/delete -> get-node-pools / *-node-pool
View JSON mirrors the console normalizers exactly (visor.ts/compute.ts/platform.ts)
so the FE renders with no change. No fabrication: GPU rows are real accelerators of
real GPU machines, clusters are real node pools, and telemetry Visor lacks is
omitted (renders — not 0). Auth: KMS service credential (Basic) or forwarded bearer.
Tests: tenant scoping/isolation, machine/gpu/cluster shape, GPU slug derivation,
launch quote+real+delete. go build ./... + go vet + go test all green.
The console Environments/Pipelines/Builds/Releases pages rendered "not
connected" because they call top-level REST that no cloud subsystem served.
Serve them natively from the platform control plane, DERIVED from the SAME
per-org project/app/deployment/build records (no new data model, no fabrication):
- GET /v1/environments — distinct Application.Environment targets across the
org's apps, each aggregating its apps (services), with a derived
type/status. List-only: an environment is a scope on apps, not a record.
- GET /v1/pipelines — one per app: its build/deploy config (repo|image) plus
the status/timing of its latest deployment. List-only: a pipeline is an app.
- GET /v1/builds — the REAL arcd BuildKit build records (platform_builds),
joined to app repo + deployment commit. List-only: builds are triggered by
the app deploy path (git source) — one trigger, not a duplicate here.
- GET /v1/releases — deployments actually applied to the cluster
(status deploying|live): a released image tag on an app/environment.
Every route is org-scoped through the same validated-principal gate (s.tenant →
requires c.User()); the response is the exact `{ "<plural>": [...] }` wrapper the
console FE normalizers read. Three org-wide store aggregates back them
(ListAllApplications / ListDeploymentsByOrg / ListBuildsByOrg), org the only
tenancy predicate. Real records or an honest empty — never fabricated history.
Tests: shape (200 + wrapper + derived fields), org isolation (second org sees
empty, no cross-tenant leak), forgeable-org refusal (no X-User-Id → 403).
Add clients/do: an org-scoped facade over digitalocean/godo's native VPCs
and LoadBalancers services, backing the console's VPC + Load Balancers pages
(which render 'not connected' today because nothing serves them).
- Routes: GET/POST /v1/vpcs, GET/DELETE /v1/vpcs/:id and the same for
/v1/load-balancers. Real godo calls, honest empty/error states, never
fabricated.
- Tenant isolation: DO is a single account, so a resource's physical DO name
is 'o'<orgHash>-<friendly> via provisioning.BucketName (the SAME org-hash
convention clients/s3 uses). List filters the account inventory to the
caller's prefix; get/delete confirm prefix ownership before acting; a
cross-tenant id reads 404 (existence-oracle guard).
- Fail-closed: absent DO_API_TOKEN every op is an honest 503.
- FE shape matches console2 VpcModule/LoadBalancerModule verbatim (vpcs[],
loadBalancers[] with the exact field names).
- Registered as subsystem 'do' (order 123). godo v1.197.0 added to go.mod.
- Tests: per-org VPC + LB isolation, forge-path 403, fail-closed 503.
Same class as c2a7534: upstream force-re-tagged luxfi/age v1.5.0 and
luxfi/pq v1.0.3 (content moved, /go.mod hashes unchanged), so the committed
zip h1: sums no longer match the served bits and `go build ./...` fails
verification. These deps entered the graph via hanzoai/commerce/metering
v0.1.2 (the agents scheduler/billing). Re-record the current zip hashes.
The console2 Agents dashboard calls two org-wide routes that were never
reachable: the bare /v1/agents/:name wildcard captured "metrics"/"activity"
as an agent name (Fiber matches in registration order), so both 404'd and the
dashboard rendered a permanent "not connected" state.
- Register the two static routes BEFORE :name so they win the match.
- GET /v1/agents/metrics?range=24H|7D|30D -> a per-agent invocations-over-time
histogram bucketed from REAL agent_runs rows; the Resource Usage rollup is
all-null because this store meters no CPU/mem/storage/cost (honest em-dash,
never a fabricated trend). Shape mirrors console2 normalizeMetrics exactly
({range,series:[{key,points:[{t,v}]}],resource:{...}}).
- GET /v1/agents/activity -> org-wide recent-activity feed: each recorded run
is an invoked/failed event, each agent's own create/update timestamps are
created/updated events; merged newest-first, capped 50. Shape mirrors
normalizeActivity ({activity:[{id,kind,agent,message,at}]}).
- Store: add RunsSince(org,since,limit) — org-wide runs across all agents,
tenancy on the org column; powers both surfaces.
- Tests prove the surfaces are not shadowed (200, not 404), reflect only real
runs, and stay org-isolated.
deps.AI was permanently nil under the default all-enabled config
(pickAIClient returned nil for cfg.Enabled("ai") and no in-process ai
subsystem ever filled it), so every agent run 503'd "inference is not
configured on this deployment" before the already-live per-org metering.
New AI client (clients/aihttp.go), two credential modes:
- AIHTTPAt: static-key OpenAI-compatible client (CLOUD_AI_API_KEY) — an
operator/pre-provisioned-key override.
- AIHTTPM2M: durable default — mints+auto-refreshes an IAM client-
credentials token from the binary's OWN identity (IAM_CLIENT_ID/SECRET)
via x/oauth2/clientcredentials. No static key to rotate, no expiry cliff,
no new secret to store. On Hanzo the identity resolves to
admin/hanzo-cloud, which the gateway treats as balance-exempt, so cloud's
per-org ResourceMeter stays the single revenue debit (no double-bill).
- build.go pickAIClient: static key -> M2M -> ZAP RPC -> fail-closed stub.
Never returns nil (the live bug). Secret never logged.
- config.go: CLOUD_AI_BASE_URL (default https://api.hanzo.ai/v1),
CLOUD_AI_API_KEY (optional), CLOUD_AI_DEFAULT_MODEL (default
deepseek-v4-flash), AIAuthClientID/Secret from IAM_CLIENT_ID/SECRET.
- clients/aihttp_test.go: httptest OpenAI emulation — default-model
substitution, content parse, 4xx/5xx mapping, empty-choices error, and
M2M token mint+use+cache.
Composes with the live metering (Gate/MeterUsage) untouched. Model routing
is the gateway's job; empty model -> cheap default is the only cloud-side
fallback (no in-code model aliasing).
Addresses Red MED-1/MED-2/INFO-3:
- MED-1: revenue.Configured now means the source was actually READ (listOrgs
succeeded), not merely wired. A transient IAM failure → configured:false, never a
fabricated zero that flips margin negative into a false 'burning' alarm. Per-org
read failures mark the commerce source not-ok (partial), never presented as whole.
- MED-2: /v1/costs is a fleet-wide god-view — commerce resolves the namespace from
COMMERCE_SERVICE_ORG, NOT from a request header (it never reads X-IAM-Org-Id).
Dropped the no-op org arg from costs(); corrected the false 'resolves namespace
from X-IAM-Org-Id' claims in commerceClient + orgSubject docs.
- Test: TestFinance_RevenueSourceDown_NoFabrication proves no fake revenue/margin
when the IAM org list is unreadable while COGS still flows.
(INFO-3 false-green fix lands console-side in financeHealth.) 10 admin tests green.
The finance board's cost side now CONSUMES commerce /v1/costs (the single
vendor-COGS source of truth: DigitalOcean compute + the LLM providers we resell)
instead of re-reading DigitalOcean's billing API to derive a DO-only cost. This
removes cloud's duplicate DO COGS read and gives the board the multi-vendor
per-vendor breakdown for free.
- commerceClient.costs() reads GET /v1/costs over the admin S2S service token
(COMMERCE_SERVICE_TOKEN, no IAM user → commerce requireCostsAdmin M2M path).
- financeCost carries {configured,totalCents,vendors[],period}; margin cost is
now the multi-vendor TotalCents, not DO month-to-date spend.
- DigitalOcean stays ONLY as the orthogonal promo-credit/runway treasury view
(commerce does not track our prepaid credit); its MTD spend feeds runway alone.
- /v1/admin/finance shape is additive (cost.digitalocean preserved) and the
global-admin guard is unchanged.
- tests: fake commerce now serves /v1/costs; margin = revenue - COGS; DO-off path
proves COGS still flows from commerce (decoupled).
A brand-new tenant's namespace is created by ensureNamespace, but the operator's
tenant-RBAC controller projects cloud-api's `cloud-api-platform` RoleBinding
(get/create resourcequotas/limitranges/services.hanzo.ai in tenant-<org>)
ASYNCHRONOUSLY. The first-ever deploy raced ahead of that RoleBinding and failed
with `resourcequotas ... is forbidden`, self-healing only on a manual retry.
Gate ensureNamespace on a SelfSubjectAccessReview readiness poll
(waitForTenantRBAC): before touching the quota objects, ask the apiserver — as
cloud-api's OWN identity — "can I get resourcequotas in tenant-<org>?" and wait,
with bounded exponential back-off (~45s ceiling), for the operator's RoleBinding
to land. An already-onboarded tenant is confirmed by a single fast probe (no
sleep), so existing deploys are not slowed. On timeout, fail CLOSED with a
retryable errTenantProvisioning (deployErrStatus -> HTTP 503, honest
"provisioning, retry") — never a fabricated success. Creating a SSAR needs no
tenant RBAC (system:basic-user), so the probe is itself immune to the window it
closes. No RBAC or namespace derivation is loosened.
Tests (clients/platform/tenant_rbac_test.go): retry-until-RoleBinding-lands
succeeds end-to-end (quota+Service CR written); bounded timeout fails closed
(503, no quota/CR written); ctx-cancel aborts promptly; ready tenant resolves in
one probe (auto-glue fast path unchanged).
Lands the agent-backend metering feature onto the zip->zap-proto-migrated main
(cloud is already fully migrated on main; 8 subsystem pins + MountAll clean).
- /v1/agents/* self-meters a per-run fee to commerce: pre-authorize the org's
prepaid credit balance fail-closed (402 insufficient_balance), debit on success,
attributed to product "agent". Added to selfMeteredPrefixes so the edge gate
never double-bills. Run path (money-moving) requires a VALIDATED principal
(c.User() non-empty), refusing the no-bearer forge path; scheduled runs carry an
unforgeable 'scheduler'-prefixed actor.
- ResourceMeter.Gate now forwards costCents as AuthInput.AmountCents so the gate
enforces available >= fee (not merely > 0) — a 1-cent balance can no longer
authorize a run that takes the ledger negative. MeterUsage generalizes the
per-org debit (Actor/Model/token attribution) while Meter keeps its signature.
- Long-running agents: cron scheduler scans once a minute over a partial index
(ix_agents_scheduled), bounded per-org cap (CLOUD_AGENT_MAX_LONG_RUNNING);
scheduler.stop drains in-flight runs inside the SIGTERM budget via ShutdownAll.
Deps: commerce/metering v0.1.0 -> v0.1.2 (Actor field + AuthInput.AmountCents).
All other pins inherited from migrated main (ai v1.789.1, authz v1.10.3,
base v1.4.6, commerce v1.42.29, licensing v0.1.1, metrics v0.4.1, o11y v1.3.12,
vfs v0.4.4). hanzoai/zip stays out of the graph; zip == zap-proto/zip v1.2.0.
go mod verify clean; go build ./... EXIT 0; MountAll boot-smoke clean (no
want *zip.App); agents -race + root/ml/provisioning billing tests green.
Verified the /api/ -> /v1/ rip is COMPLETE on origin/main: zero owned /api/
route registrations, zero owned /api/ client strings. The rip landed in
cf58f8d (productsvc: drop residual /api/ prefix) plus the o11y and eval
cleanups. Every remaining /api/ reference is a non-owned external contract:
- clients/o11y/o11y.go, o11y_test.go: doc comments ("no /api/, no rewrite")
documenting the now-removed upstream rewrite.
- clients/pricing/pricing.go: https://openrouter.ai/api/v1/models — a
third-party vendor URL.
- clients/platform/*_test.go: "apps/api/deploy" where `api` is a user's
APP NAME inside /v1/platform/... paths (not an API prefix).
- zapface/dispatch.go: comment already says "the /v1 convention".
- clients/prompts/catalog.json: a prompt-catalog data blob describing a
different project (prompts.chat), not this repo's routes.
The IAM /api/add-usage-record callout referenced in older cloud docs is a
Casdoor/casibase-lineage endpoint the LEGACY Node cloud-api used (documented
in hanzoai/commerce auth/iam_admin.go). The current Go cloud-api does NOT
call it: billing meters to commerce via hanzoai/commerce/metering. No
cross-service IAM usage-record dependency exists in this repo.
The only change here is a deps-hygiene fix so the build verifies clean:
luxfi force-re-tagged age@v1.5.0 and pq@v1.0.3, drifting their module-zip
h1: hashes vs the recorded go.sum (go.mod hashes unchanged). Realigned to
the upstream hashes at the SAME versions — no major/minor bump.
CGO_ENABLED=0 GOWORK=off go build -mod=readonly ./... -> exit 0
go test ./clients/{o11y,eval,pricing,ml,admin} ./zapface . ./clients/platform -> all ok
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
applyService (operator Service CR write) was not jointly ordered with
FinalizeLive (live-pointer DB write): under concurrent same-app deploys an
OLDER deploy's CR write could land AFTER a NEWER one already went live, leaving
the live Service CR image lagging the recorded live version. deployImage had no
supersede gate at all.
Introduce applyLive — the ONE deploy mechanic shared by the image-source path
and the git build reconciler — running supersede-check → applyService →
FinalizeLive as one per-app-serialized critical section (appMutex, fixed-shard,
O(1) memory). An older deploy that loses the race is superseded and never writes
its CR. FinalizeLive's monotonic CAS still backstops DB monotonicity.
go vet + full clients/platform test suite green.
HIGH-2 — cloud-api touches NO K8s Secret:
- delete ensurePullSecret + secretsGVR + CLOUD_PLATFORM_PULL_DOCKERCONFIG;
the per-tenant ghcr-pull Secret is provisioned by the OPERATOR's tenant-RBAC
controller (from a KMS-synced source). serviceCR only REFERENCES it by name.
cloud-api's ServiceAccount holds no secrets grant and issues no Secrets call.
MED-1 — monotonic live-version finalize (no build-time inversion):
- Store.FinalizeLive: ONE atomic conditional UPDATE that advances an app to live
ONLY when its version >= the currently-live version (no read-then-write TOCTOU).
- reconcileBuild gates on buildSuperseded before applying the (older) CR, and
records a late/older build 'superseded' (image still succeeded) instead of
regressing the running workload. deployImage shares the same FinalizeLive.
- tests: TestFinalizeLiveIsMonotonic (store CAS) + TestBuildReconcilerVersionMonotonic
(e2e inversion: newer-first-live, older-late-superseded, CR never downgraded).
The git deploy path launched a BuildKit Job fire-and-forget and left the
deployment stuck 'building' — deploy.go documented the build watcher as
'phase 2'. Implement it as the ONE owner of the handoff:
- reconcile.go: a restart-safe reconciler (state in the store, not a
goroutine) that scans 'building' deployments, checks each build Job, and on
success applies the operator Service CR with the built image (the SAME
applyService the image path uses) → deployment 'deploying', app 'live'. On
failure/deadline it records the honest error. Started from Mount, stopped on
Shutdown. Org-scoped: every write targets tenant-<row.Org>.
- store.ListBuildingDeployments: cross-org 'building' query (reconciler input).
- k8s.jobOutcome/jobResult: ONE Job terminal-state classifier shared by the
concurrent-build cap and the reconciler.
- k8s.serviceCR imagePullSecrets + ensurePullSecret: the built image is PRIVATE
(ghcr.io/hanzoai/tenant-<org>/*); provision the tenant GHCR pull secret from
CLOUD_PLATFORM_PULL_DOCKERCONFIG (KMS-synced; no-op when unset) and reference
it so the operator's pod can pull.
Tests: jobOutcome classifier + ListBuildingDeployments (oldest-first, cross-org,
building-only). Full platform suite green (RED authz/cmd-injection incl.).
clients/eval/telemetry.go opened a SECOND direct clickhouse.Open with a
parallel CLOUD_EVALS_CLICKHOUSE_* cred namespace, bypassing the shared ZAP
datastore mesh that clients/analytics + ai/object already use. Consolidate:
- eval telemetry now routes every write/read over ai/object's shared client
(aiobject.DatastoreExec / DatastoreQuery / DatastoreEnabled), the same peer
the o11y ledger + /v1/analytics use. One connection, one pool, one
retry/backoff, one KMS-injected cred namespace (DATASTORE_*).
- DELETE the CLOUD_EVALS_CLICKHOUSE_* namespace and the private clickhouse.Open.
- Ownership stays clean: eval owns only its two tables (hanzo.eval_traces,
hanzo.eval_scores); ai/object owns hanzo.cloud_usage / hanzo.observations.
- No batch primitive needed — eval Records are single-row, mapping to the
shared DatastoreExec INSERT ... VALUES (?) the o11y write path already uses.
- Async-connect aware: readiness gates per-op on DatastoreEnabled() (honest
'unavailable' in the boot window), tables ensured idempotently, latched once.
- Reads bind org + narrowers positionally (?) — no interpolation; LIMIT always
applied. Tenant isolation + score finiteness invariants unchanged.
provisioning's direct CH client is a distinct control-plane concern — untouched.
go build + go test ./clients/eval/... green.
A green go build/vet/test does not catch a binary that PANICS at startup.
v1.786.14/.15/.16 compiled clean but crashed at boot with
cloud: mount metrics: metrics.Mount: app is *zip.App, want *zip.App
(a runtime type-assert from an incomplete hanzoai/zip -> zap-proto/zip
migration), published green, and CrashLooped in prod. The only gate that
catches this class is running the binary.
Restructure the single build-push into build(load) -> smoke -> push:
1. Build once to a local cloud:smoke tag (push:false, load:true), warming
the BuildKit builder cache.
2. Boot that exact image with a minimal, prod-representative env (writable
ephemeral /data + a throwaway 32-byte KMS master key so the KMS plane
mounts on its normal ready path) and assert it reaches "listening" with
NO startup-crash signature (metrics.Mount / mount metrics / panic /
want *zip.App), else exit 1 BEFORE any push. Container always rm -f.
3. Re-run build with push:true and the real tags: identical context /
platform / secrets, so every layer is a cache hit from step 1 and it
only publishes the already-tested image.
Stays on the self-hosted arcd amd64 scale set + GH_PAT; notify-universe
unchanged.
Proven to DISCRIMINATE against the published images:
ghcr.io/hanzoai/cloud:v1.786.18 (known-good) -> SMOKE PASS (exit 0)
ghcr.io/hanzoai/cloud:v1.786.15 (known-bad) -> SMOKE FAIL (exit 1)
SECURITY: cloud origin/main (55b073e7) had DIVERGED from the LIVE image
v1.786.18 (faef12aa) at merge-base 862d4623 — v1.786.18 was tagged+deployed
from fix/cloud-1.786.18-zip-and-gates but NEVER merged back to main. It carries
the F1 forged-X-Org-Id cross-tenant gate that main LACKED:
- clients/principal (validated-principal helper) — new package
- bot + o11y reverse-proxy forge gates (bot.go, o11y.go + red_forge_test.go)
- whole-data-plane principal gating across agents/crm/eval/functions/git/kms/
ml/plan/pricing/projectsvc/prompts/provisioning/s3
Building .20 from main+platform ALONE would have REGRESSED this live fix
(reopened the forged-X-Org-Id hole — the ".17 insecure, do NOT deploy" hole).
This merge makes the artifact a true SUPERSET of the security floor.
Only conflict: clients/provisioning/provisioning.go import block — resolved to
keep BOTH the F1 principal.Validated(c) gate AND the platform sanitizeOrg
injectivity crit-fix (OrgHasUnsafeRune + raw-byte hash). go.mod/go.sum merged
clean (main and .18 converged on identical dep floors: ai 1.789.1, authz
1.10.3, base 1.4.6, commerce 1.42.29, o11y 1.3.12, vfs 0.4.4, zap-proto/zip).
Result = main (analytics/crm/templates/git + zip migration) + /v1/platform mount
+ v1.786.18 F1 gate. All 24 subsystems registered. go build ./... + go vet green.
Tests green together: platform (TestRED_CommandInjectionBlocked, cross-tenant,
injective), F1 (TestRed_BotProxyForwardsForgedOrgNoPrincipal,
TestRed_O11yProxyGatesForgedOrgNoPrincipal), provisioning, principal, root cloud.
Merge RED-PASSED blue/paas-v1platform@86cb4c15 onto main@55b073e7, mounting the
per-org container-app PaaS control plane at /v1/platform (HIP-0106) — the deploy
engine behind one-click app deploys (ERP/Helpdesk ride on it).
Conflicts resolved (5 files, all combine-both-sides — no logic dropped):
- subsystems/subsystems.go: KEEP every registration — platform (order 124) +
analytics/crm/git/templates/prompts/agents/functions + all pre-existing.
- clients/provisioning/provisioning.go: main's zip canonical import +
branch's sanitizeOrg injectivity crit-fix (OrgHasUnsafeRune reject,
raw-byte SHA-256, no TrimSpace) — both preserved.
- middleware_identity.go: main's zap-proto/zip + doc; branch's OrgHasUnsafeRune
(root cloud pkg) + SanitizeIdentity handler hardening — both preserved.
- provisioning_test.go / middleware_identity_test.go: both test sets kept.
Integration: main migrated the repo hanzoai/zip -> zap-proto/zip; the branch
predated it, so the new clients/platform/{deploy,platform,http_test}.go were
rewritten to the canonical github.com/zap-proto/zip (incompatible zip.Ctx types
otherwise). go.mod unchanged from main (zap-proto/zip v1.2.0 direct); no new dep
(k8s.io/{apimachinery,client-go} v0.35.0 already present via ml/paassvc).
Crit-fixes preserved EXACTLY (RED-PASSED, proven green):
argv build (TestRED_CommandInjectionBlocked), org-slug injective
(TestSanitizeOrg{Injective,WhitespaceInjective}, TestBuildImageRefIsInjective),
ResourceQuota (TestEnsureNamespaceAppliesQuota), no cross-tenant
(TestHTTPCrossTenantIsolation, TestNamespaceIsDerivedFromOrgNotInput,
TestServiceCRAlwaysPinnedToTenantNamespace).
go build ./... + go vet + go test (platform/provisioning/root cloud) all green.
Startup crash on v1.786.15/.16: `cloud: mount metrics: metrics.Mount: app is
*zip.App, want *zip.App`. Cloud's core migrated to github.com/zap-proto/zip
(v1.786.13→.15) so app is *zap-proto/zip.App, but eight hanzo modules cloud
imports still pinned the OLD github.com/hanzoai/zip and registered subsystems
that type-assert app.(*hanzoai/zip.App). Distinct import paths = distinct Go
types, so MountAll's first such subsystem ("metrics", hanzoai/metrics@v0.4.0)
failed the assert at runtime (build/vet/test stayed green — the mismatch is
runtime-only). authz/base/o11y were the same latent break behind it.
Forward fix (the canonical-home migration was already released upstream; cloud
merely lagged): bump every lagging module to its migrated tag —
ai v1.789.1-… → v1.789.1 authz v1.10.1 → v1.10.3
base v1.4.1 → v1.4.6 commerce v1.42.27 → v1.42.29
licensing v0.1.0 → v0.1.1 metrics v0.4.0 → v0.4.1
o11y v1.3.7 → v1.3.12 vfs v0.4.1 → v0.4.4
authz pinned to v1.10.3 specifically: v1.10.2/v1.10.4 carry an unrelated
GetPolicy 2-value change that breaks the pinned hanzoai/iam; v1.10.3 has the zip
migration AND the iam-compatible 1-value signature. commerce pinned to v1.42.29
(v1.43.0 regressed back to hanzoai/zip). clients/analytics (from the merged .16
work) is cloud's own code and is migrated in-place. Result: hanzoai/zip is gone
from go.mod/go.sum and the whole module graph; the compiled binary mounts
metrics+o11y+all subsystems and reaches "listening" with no panic.
Also folds in the COMPLETE F1 close (RED found the gate was partial — two
reverse-proxy paths still forwarded a forged X-Org-Id):
- clients/bot: gate proxy() on principal.Validated before forwarding X-Org-Id to
bot-gateway (RED PoC red_forge_test.go now passes: no-principal forge → 403).
- clients/o11y: wrap the installed reverse-proxy handler in gate() — refuse any
request with no X-User-Id before it reaches the o11y runtime (forge twin test).
- clients/crm: extend the forge guard to WRITE+DELETE verbs (belt-and-suspenders).
- middleware_identity: refresh the stale FAIL-MODE comment — post-F1 the DATA
plane also fails secure on a cold-cache JWKS failure (bounded by stale-on-error).
Base = F1 (fix/cloud-data-plane-principal-gate, bff2a688) + origin/main (analytics
.16, 862d4623). One healthy image: F1 + analytics + zip-fix + bot/o11y gates.
Aligns clients/analytics with serve.go + every other in-tree clients/* subsystem,
which import github.com/zap-proto/zip. The original commit imported the OLD
github.com/hanzoai/zip, making analytics.Mount assert *hanzoai/zip.App while serve
passes *zap-proto/zip.App — a boot-time mount type mismatch. Removes the last in-tree
hanzoai/zip importer so once the migration lane re-releases the external subsystem
modules on zap-proto/zip, main boots with ONE zip.App type. (The DEPLOYED v1.786.17
is built off v1.786.13, which is all-hanzoai/zip, and is unaffected.)
Adds clients/analytics (order 132, registered as analyticssvc so the real
/v1/analytics/health owns the probe, not serve.go's generic liveness route) —
the backend for the console Native Analytics module (unified-analytics.md §5).
Two read lenses over the ONE hanzo warehouse, reusing the SAME clickhouse-go/v2
client the ai o11y ledger opens (ai/object DatastoreQuery/DatastoreEnabled/
EnsureCloudUsageTable/ResolveCloudUsageWindow) — no second CH client, DRY:
- LLM lens (REAL): hanzo.cloud_usage — requests/tokens/spend/models/errorRate
- web+commerce lens: hanzo.events — honest-empty until the collector emits
Surface (read-only, org-scoped, /v1):
GET /v1/analytics/overview per-org KPIs (llm real; web/commerce honest-empty)
GET /v1/analytics/timeseries requests/tokens/spend over hour|day buckets
GET /v1/analytics/top top models (real) + top products (honest-empty)
GET /v1/analytics/health datastore connectivity + lens-table availability
Tenant isolation is the security bar: tenant() requires a VALIDATED principal
(c.User(), set by SanitizeIdentity only for a verified bearer) AND a valid org
(c.Org(), the minted owner claim) — closing the Phase-1 no-bearer forged-X-Org-Id
data path exactly as clients/s3 does. Every query binds the org POSITIONALLY
(query.go llmWhere/eventsWhere), so a maxpower token can never read another org.
ClickHouse creds are KMS-injected env (DATASTORE_*), never hardcoded.
Tests: query-boundary isolation (org bound, never interpolated, incl SQLi slug),
honest-empty, real-number KPIs, errorRate, gap-filled series, top-models pct;
HTTP: no-principal->403, forged-org-no-bearer->403, datastore-down->honest 503,
bad-range->400, health owned-by-analytics honest 503 when down.
The org identifier was TrimSpace'd at both trust-boundary sites
(middleware_identity.go on claims.Owner + client X-Org-Id, and
provisioning.sanitizeOrg before hashing), so two DISTINCT IAM orgs differing
only by edge/internal/unicode whitespace ('acme' vs 'acme ' vs 'ac me' vs an
NBSP/ZWSP variant) collapsed onto ONE tenant-<slug> namespace / image ref /
bucket / DB — a cross-tenant fold (IAM org name is an unvalidated varchar, so a
fold-sibling is registrable and mints a valid token).
FIX — normalize+VALIDATE at the trust boundary, reject rather than fold:
- cloud.OrgHasUnsafeRune: refuse any org bearing a whitespace / control /
zero-width-format (Cf) rune. fasthttp OWS-trims header values, so folding
such an org could never round-trip through transport — rejection (fail
secure) is the only injective option. Visible case/'.'/'-' still fold
injectively via the org-slug hash.
- middleware_identity.go: owner is taken verbatim from the validated principal
(no TrimSpace) and refused if unsafe -> request resolves org-less, every
tenant() gate fails closed 403. Client X-Org-Id refused likewise.
- provisioning.sanitizeOrg: reject unsafe-rune inputs (defense-in-depth for
non-header callers e.g. clients/s3) and hash the RAW bytes, never a trimmed
copy. c.Org() is now the sole tenancy source and injective end-to-end.
Regression tests: {acme, 'acme ', 'ac me', NBSP/ZWSP/BOM/tab variants} ->
distinct-or-rejected, never colliding (provisioning + platform + middleware
JWT-owner path). go build ./... + go vet + affected suites green.
/v1/platform (PaaS) — RED do-not-ship findings. Tenancy core untouched.
CRIT-1 — OS command injection in the privileged BuildKit Job:
launchBuildJob now emits buildctl as EXEC-FORM argv ([]string, no `sh -c`),
so no shell parses any input. repo.url / dockerfile / git-ref are validated
(validate.go): https-only URL to an allowlisted git host, no shell/flag
metachars; safe relative dockerfile (no `..`); safe branch/tag/commit ref
(no `#`, no metachars). Output image ref is forced server-side — a client
cannot override --output/--opt. Validation runs at the build choke point AND
early at createApp (400). red_cmdinj_poc_test.go flipped to a passing guard.
CRIT-2 — sanitizeOrg collision (non-injective) → cross-tenant takeover:
deleted the lossy platform.sanitizeOrg; tenant()/tenantNamespace()/
buildImageRef() now use the ONE injective provisioning.SanitizeOrg (DRY,
reused — commit 4e77020c). Image ref made injective too: org+app are now
separate '/'-joined path components (ghcr.io/hanzoai/tenant-<org>/<app>),
neither slug can contain '/', so (a-b,c) vs (a,b-c) no longer collide.
MED-3 — quotas / replica bounds / shared-build DoS:
clampReplicas caps replicas to [1,20] (env CLOUD_PLATFORM_MAX_REPLICAS) at
createApp, applyService, and scaleService (fail-secure default). ensureNamespace
applies a ResourceQuota + LimitRange per tenant namespace (idempotent).
launchBuildJob caps concurrent builds per org (default 3, errTooManyBuilds→429).
Tests: cmd-injection blocked (5 vectors), org-slug + image-ref injectivity,
replica clamp (unit+HTTP), namespace quota/limitrange, concurrent-build cap.
go build ./... green; clients/platform + provisioning + s3 tests green.
A custom domains[] entry was rendered straight into the operator Service CR
ingress.hosts, so a tenant could claim another org's host or a Hanzo apex
(api.hanzo.ai) and the operator would serve an Ingress for it. Require every
custom host to be under the caller's OWN '<org>.<sitesHost>' subtree (e.g.
maxpower may only claim *.maxpower.hanzo.app); anything else is refused 501
(verified arbitrary custom domains are phase-2 domain CRUD). Closes the
cross-tenant/apex domain-hijack vector reachable in the first slice. +2 tests
(unit + HTTP); 23 tests green.
/v1/platform mutates cluster state (operator Service CRs + BuildKit Jobs in
tenant-<org>) — more consequential than a data read — so trusting X-Org-Id alone
would let a direct-to-pod caller forge X-Org-Id:victim with NO bearer and
deploy/read into another tenant (SanitizeIdentity's documented Phase-1 residual).
tenant() now gates on c.User() (X-User-Id, set ONLY for a validated principal),
mirroring the Red-hardened clients/s3.tenant. Proven live: forged X-Org-Id with
no token -> 403 (was 200); real JWT -> 200; forged header + real token ->
validated owner wins. Every legitimate caller (gateway/console BFF) carries a
user-bound bearer, so no real client breaks.
Port the standalone Dokploy (platform.hanzo.ai) tRPC backend into the unified
cloud binary as clients/platform, mounted at /v1/platform (HIP-0106). Per-org,
IAM-validated, Base/SQLite store; the deploy path writes an operator hanzo.ai/v1
Service CR into the caller's OWN tenant-<org> namespace (derived from the
validated X-Org-Id, never a request input) and the operator reconciles it. Git
apps build via an in-cluster BuildKit Job (arcd model); image apps deploy
directly. Complements clients/paassvc (admin fleet board) and clients/projectsvc
(static sites) with the container-app PaaS.
Goa is the design-first contract (clients/platform/design, goa gen -> OpenAPI 3);
the runtime is native zip handlers (one router, behind SanitizeIdentity) — the
generated net/http server is deliberately NOT mounted so the identity trust
boundary is not routed through the fiber<->net/http adaptor.
- store.go: projects/applications/deployments/builds, org column tenancy
- k8s.go: tenant-<org> namespace derivation + Service CR apply/scale/delete + BuildKit Job
- platform.go: Mount + project/app CRUD + tenant() gate + health
- deploy.go: deploy/start/stop + deployment history/logs, fail-closed (no fabricated success)
- 20 tests: store CRUD, cross-tenant isolation (RED bar), fail-closed deploy,
fake-cluster deploy-into-tenant-ns success, secret-env rejection
go build ./... green; go test ./clients/platform/ green.
2026-07-01 16:20:48 -07:00
1366 changed files with 255852 additions and 8759 deletions
| grep -v '.github/workflows/containment.yml:'; then
echo "::error::found a build/release invocation passing -tags controlplane — clients/controlplane's stub crypto must never enter a release/serve binary (see clients/controlplane/doc.go)"
| grep -v '.github/workflows/containment.yml:'; then
echo "::error::found a reference to testing.testBinary outside the Go toolchain itself — this is the linker var that spoofs testing.Testing() in a real (non go-test) binary; the containment.go runtime guard trusts that signal, so setting it anywhere in a real build path defeats it (see doc.go)"
hits=1
fi
if [ "$hits" -ne 0 ]; then exit 1; fi
echo "OK: no build/release path sets -tags controlplane or spoofs testing.Testing()"
- uses:actions/setup-go@v5
with:
go-version-file:go.mod
- name:go env for private modules
env:
GH_PAT:${{ secrets.GH_PAT }}
# GOPRIVATE names exactly the namespace that is private. github.com/hanzoai/*
# is: ai, account, commerce, orm, xorm, beego, csqlite and ~30 more are
# private repos, so they must resolve direct+authenticated and skip a sumdb
# that cannot see them. Everything else stays on the public proxy + checksum
# db, which is what makes a module hash immutable: zap-proto (all 55 repos)
# and luxfi (all 37 deps here) are public and proxy-served.
#
# This previously named zap-proto — public, and never the reason anything
# here was direct — and then set GOSUMDB=off to compensate for hanzoai/*
# being absent, which disabled checksum verification for EVERY module in the
# build, public ones included. Naming the private namespace is what the off
# semver-correct compare: the lowest of {V, FLOOR} under `sort -V` must be
# the FLOOR, i.e. V >= FLOOR. (sort -V orders v1.4.2 above v1.4.10 too.)
low="$(printf '%s\n%s\n' "$V" "$FLOOR" | sort -V | head -1)"
if [ "$low" != "$FLOOR" ]; then
echo "::error::github.com/hanzoai/zen is pinned at ${V}, below the streaming-fix floor ${FLOOR} — this re-breaks SSE streaming (empty completions). Re-pin zen to >= ${FLOOR} in go.mod before merging."
exit 1
fi
echo "OK: zen ${V} is at or above the streaming-fix floor ${FLOOR}"
- name:positive proof — clients/controlplane is unreachable from the default build
run:|
set -euo pipefail
go build ./...
for m in $(go list ./cmd/...); do
if go list -deps "$m" | grep -qx 'github.com/hanzoai/cloud/clients/controlplane'; then
echo "::error::$m links clients/controlplane into a real binary — containment breach"
[ -n "$d" ] || { echo "::error::decomplection artifact ghcr.io/hanzoai/${repo}:latest is not published — refusing to cut a release that would embed a stale/placeholder ${repo}"; return 1; }
echo "MIGRATION SMOKE INFRA FAIL: baseline $BASELINE did not reach \"listening\" — cannot stage the prior schema (inspect/bump SMOKE_MIGRATION_BASELINE)"
exit 1
fi
docker stop "$B1" >/dev/null
chmod_vol
# Boot 2 — the candidate migrates that on-disk schema IN PLACE. This is the gate.
2. **Agent-NAME axis** — needs (1): the agent run debit records `provider=agent` +
`model=<llm>` but not the agent name, so `metadata.agent` stays honest-empty
until commerce persists an `agent` field the agents meter sets to `a.Name`.
3. **compute split** — `ml` (predict) and `visor` (GPU) both meter `provider=compute`;
read-side can't split `inference` vs `gpus`. Needs (1) so each sets its product id.
4. **exec / containers** (`clients/exec`, Code Interpreter) — authed by a shared
service key (X-API-Key), NO per-org identity, so it can't meter per-org; its
compute is billed upstream at the chat/agent layer that invokes it.
5. **playground** — routes to `/v1/ai/*`, already metered as AI inference.
## 5. Secrets
- Per-tenant, KMS-managed only (`kms.hanzo.ai`, KMSSecret CRDs). No shared
service key stands in for a user. The only service tokens that exist are
narrow, per-tenant, and never used to impersonate a user for LLM spend.
- A surface's own OIDC registration is a PUBLIC client — there is no client
secret to store.
## Surface conformance (as of this contract)
| Surface | Login (PKCE public) | Forwards user token | Org from `owner` | Billed via cloud (org,project) |
|---|---|---|---|---|
| **studio.hanzo.ai** | ✅ reference | ✅ (validates locally) | ✅ | ⚠️ renders run on studio's own GPU workers and self-report to commerce keyed by org via a per-tenant commerce token — org-keyed, but not the forward-bearer-to-gateway path (studio does not call the cloud LLM gateway for its core renders) |
| **console.hanzo.ai** | ✅ | ✅ same-origin `/v1` through the gateway | ✅ | ✅ (it IS the canonical consumer) |
| **hanzo.app** | ❌ confidential client (`IAM_CLIENT_SECRET`, userinfo/introspect) | ✅ to its own backend; org from token `owner` | ✅ | ❌ builder AI runs on OpenRouter with an apiKey (`lib/llm/generation-api.ts`), NOT the cloud gateway — off the unified meter |
hanzo.app is the remaining gap: it needs the same treatment chat just got — switch
its IAM registration to a PKCE public client, and route its builder AI generation
through api.hanzo.ai forwarding the user's IAM bearer so usage meters against the
MODERNC="$(CGO_ENABLED=1 go list -tags "libsqlite3 sqlite_fts5" -deps ./cmd/cloud 2>/dev/null | grep -c 'modernc.org/sqlite'||true)";\
["$MODERNC"="0"]||{echo"SQLITE-GATE FAIL: cmd/cloud links modernc.org/sqlite ($MODERNC pkgs) under CGO=1 — double-registers \"sqlite\" with hanzoai/sqlite(mattn) and panics at init.";exit 1;}
# RED gate — ENCRYPTION PROOF + the cek.go GOLDEN-VECTOR KAT, under the SAME CGO +
# libsqlcipher build this image ships. TestEncryptionProof asserts real
# ciphertext-at-rest (SQLITE_REQUIRE_CODEC=1 makes a plaintext link FAIL → NO
# image). TestUnwrapGoldenFixture asserts a FROZEN pre-luxfi-swap 61-byte DEK
# sidecar still decrypts under the shipped luxfi/crypto-AEAD code — existing
# encrypted stores stay readable, or NO image.
RUN --mount=type=cache,id=cloud-gomod-v4,target=/go/pkg/mod,sharing=locked \
@echo ">> embedded real console bundle into webui/dist (index.html $$(wc -c < webui/dist/index.html) bytes)"
deploy-ui:## Build the monochrome ArgoCD dashboard bundle into clients/deploy/webui/dist (go:embed source). DEPLOY_DIR=<path to hanzoai/deploy>.
@command -v yarn >/dev/null 2>&1||{echo"yarn is required to build the deploy dashboard bundle";exit 1;}
@test -f "$(DEPLOY_DIR)/ui/package.json"||{echo"deploy checkout not found at $(DEPLOY_DIR) — set DEPLOY_DIR=<path to hanzoai/deploy on rebrand/hanzo-monochrome>";exit 1;}
agentskills:## Regenerate the FULL agent-skills catalog into clients/agentskills/catalog (go:embed source) from the openapi SOT. OPENAPI_DIR=<path to openapi>.
@test -f "$(OPENAPI_DIR)/skills.py"||{echo"openapi checkout not found at $(OPENAPI_DIR) — set OPENAPI_DIR=<path> or clone hanzoai/openapi";exit 1;}
# skills.py rewrites the whole catalog dir; the .gitignore keeps only the tiny
# `ai` fallback tracked, so the full set is embedded at build but never committed.
build-standalone:webuibuild## Build the REAL 1-binary console: console build:embed → webui/dist → go build.
hanzo:## Build the hanzo control-plane CLI into ./bin/hanzo (pure Go, same mode as cmd/cloud — registers the ONE "sqlite" driver exactly once; a plain CGO_ENABLED=1 `go build ./cmd/hanzo` links the fork's mattn backend alongside the embedded modernc importers and panics, see header).
@@ -23,11 +74,14 @@ run: build ## Run with iam,base,kms,gateway,o11y enabled (matches README quickst
smoke:## Build and run cmd/cloud-smoke (mount-time integration check).
$(GO) run ./cmd/cloud-smoke
test:## Run unit + integration tests.
$(GO)test ./...
test:## Run unit + integration tests (pure-Go, exactly as prod ships).
CGO_ENABLED=$(CGO_ENABLED)$(GO)test ./...
test-cgo:## Prove the cgo build works too — forces the fork's pure-Go backend via -tags sqlite_purego so the embedded modernc importers don't double-register "sqlite".
CGO_ENABLED=1$(GO)test -tags sqlite_purego ./...
vet:## go vet across the module.
$(GO) vet ./...
CGO_ENABLED=$(CGO_ENABLED)$(GO) vet ./...
tidy:## go mod tidy + verify go.sum.
$(GO) mod tidy
@@ -41,3 +95,6 @@ docker-push: docker ## Push the Docker image to ghcr.io. Requires docker login.
clean:## Remove built artifacts.
rm -rf bin
native:## Build the native flags evaluator staticlib (required for CGO=1 builds/tests).
@@ -15,7 +15,7 @@ docker run -p 8080:8080 ghcr.io/hanzoai/cloud:latest
## What this is
`hanzoai/cloud` is one Go binary that mounts every Hanzo subsystem (iam, kms, base, gateway, ai, commerce, vfs, mq, dns, amqp, mcp, o11y, ...) into a single multi-tenant process. Same artifact serves `api.hanzo.ai`, `api.osage.cloud`, `api.lux.cloud`, `api.zoo.cloud`, and every white-label reseller. Brand, enabled subsystems, and tenant scope are deployment configuration.
`hanzoai/cloud` is one Go binary that mounts every Hanzo subsystem (iam, kms, base, gateway, ai, commerce, vfs, mq, dns, amqp, mcp, o11y, ...) into a single multi-org process. Same artifact serves `api.hanzo.ai`, `api.osage.cloud`, `api.lux.cloud`, `api.zoo.cloud`, and every white-label reseller. Brand, enabled subsystems, and org scope are deployment configuration.
t.Fatalf("GET /v1/store/current hit the AI balance gate (got %d, body %s) — /v1/store must be a commercePrefix so it reaches commerce, not the /v1/* catch-all",code,body)
}
// And a sibling store path (the listing upsert the publish edge writes) is owned too.
t.Errorf("framework content module %q not registered — a blank import in apps.go is missing (erp drop = ledger hooks gone from the binary)",want)
}
}
// The module registry proves DocTypes are linked, but erp's ledger-posting
// HOOKS register in a separate init() step; assert them directly so the guard
// survives a future split of registerHooks() out of erp's module init().
ifframework.RegisteredHookCount()==0{
t.Error("no framework lifecycle hooks registered — erp's ledger-posting hooks (computeJournalTotals, journalEntry/paymentEntry submit+cancel, …) are not linked into the binary")
// (@150 catch-all) and the hanzoai/o11y module wildcard (@70), which main currently
// DROPS and this PR restores. The module wildcard (order 70) is now folded in as the
// terminal sub-mount of the in-repo o11y read plane (order 69), so "o11y" is ONE spec.
varfrozen=[]struct{
namestring
ownsHealthbool
hasShutdownbool
}{
{"pubsub",false,true},// was order 5
{"kafka",false,true},// was order 6
{"agentskills",false,false},// was order 8
{"flags",true,true},// was order 9; native engine: /v1/flags health + store shutdown
{"kms",true,false},// was order 10
{"metrics",false,false},// was order 40
{"ingress",false,true},// was order 42
{"account",false,false},// was order 48
{"iam",false,false},// was order 50
{"base",true,true},// was order 60; per-org embed added Shutdown (#298)
{"o11y",false,true},// ONE observability subsystem (was co-owned orders 69+70): read plane + the hanzoai/o11y module wildcard folded in as MountO11y's terminal sub-mount. OwnsHealth=false keeps /v1/o11y/health the generic always-ok route the module co-entry used to trigger.
{"authz",false,false},// was order 70
{"commerce",false,false},// was order 100
{"licensing",false,false},// was order 110
{"plan",true,false},// was order 111; enable id normalized plans->plan (routes stay /v1/plans/*)
{"pricing",true,false},// was order 112
{"storage",true,false},// was order 118
{"provisioning",false,false},// was order 120
{"billing",false,false},// was order 121
{"account-bridge",false,false},// was order 122
{"do",false,false},// was order 123
{"platform",true,false},// was order 124
{"projects",false,false},// was order 125
{"dns",false,false},// new: /v1/dns zone plane (after projects)
{"prompts",false,false},// was order 126
{"agents",false,true},// was order 127
{"link",false,true},// new: unified AI login manager (/v1/links), after agents
t.Error("gate ADMITTED a $17.376 request against a $5.00 balance — the estimate is reaching AuthorizeVerdict as 0 cents, so the size check never runs")
}
// The same balance must still admit a request it can actually cover, or the
{zen:"zen5-pro",carrier:"claude-opus-4-8[1m]",env:"ANTHROPIC_DEFAULT_OPUS_MODEL",name:"Zen5 Pro",desc:"Hanzo Zen5 Pro — DeepSeek-V4 class (1M context)"},
{zen:"zen5-pro",carrier:"claude-fable-5[1m]",env:"ANTHROPIC_DEFAULT_FABLE_MODEL",name:"Zen5 Pro (max effort)",desc:"Hanzo Zen5 Pro, top tier (1M context)"},
// api.hanzo.ai exposes the standard OpenAI /v1/models shape,
// not Codex's private remote model-catalog schema. Skip that
// optional refresh and supply the coding model's metadata here.
"-c",`features.remote_models=false`,
"-c",`model_context_window=262144`,
"-c",`model_auto_compact_token_limit=235929`,
}
},
install:install,
}
}
// zenIdentityPrompt is appended to Claude Code's base system prompt so a model
// served through the Hanzo cloud self-identifies as a Hanzo Zen model. It is an
// APPEND, not a replace: Claude Code keeps its base prompt (tool-use, safety,
// coding conventions); only the model's identity is Hanzo Zen. The model is
// served as a zen5 alias via api.hanzo.ai. Passed as --append-system-prompt,
// present in --safe too (identity is not a permission bypass).
constzenIdentityPrompt="You are running through the Hanzo AI cloud as a Hanzo Zen model (the `zen5` capability tier, served via api.hanzo.ai). When asked what model or assistant you are, identify as a Hanzo Zen model. You are operating inside the Claude Code harness; keep its tool-use, safety, and coding conventions — only your identity is Hanzo Zen."
returnfmt.Errorf("no platform token: pass --platform-token, set HANZO_PLATFORM_TOKEN, or run `hanzo login --platform-token <tok>`")
returnfmt.Errorf("not authenticated: run `hanzo login` (an IAM login now authorizes the platform; a --platform-token / HANZO_PLATFORM_TOKEN still works for machine automation)")
returnnil,fmt.Errorf("no build token: set HANZO_BUILD_TOKEN / PLATFORM_BUILD_CALLBACK_TOKEN or `hanzo login --build-token <tok>`")
returnnil,fmt.Errorf("not authenticated: run `hanzo login` (an IAM login now authorizes builds; HANZO_BUILD_TOKEN / --build-token still works for machine automation)")
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.