PAI is the privileged internal API plane. It is not publicly reachable and is protected by:
X-PAI-Key)X-PAI-Token) / session-JWT second factor (PAI_REQUIRE_PAIRED_TOKEN, PAI_SESSION_TOKENS_ENABLED)PAI_IP_ALLOWLIST) — skipped when empty; restrict at the network/system layer (WAF / load balancer / security group) instead/internal/v1
Internal Swagger UI (requires a PAI key, or a session JWT when PAI_SESSION_TOKENS_ENABLED is on; IP allowlist enforced only when configured):
GET /internal/docsGET /internal/openapi.jsonProvide a PAI key via one of:
X-PAI-Key: phx_pai_...X-API-Key: phx_pai_...Authorization: PAIKey phx_pai_...POST /api/v1/admin/pai/keys
Generated Global Admin only, no email step. Returns the api key (phx_pai_...) and companion token (phx_pait_...) once — both shown on screen; the companion token is also emailed best-effort. Manage with GET /api/v1/admin/pai/keys (list, no secrets) and POST /api/v1/admin/pai/keys/{key_id}/revoke (revoke). To rotate, generate a new pair and revoke the old key. POST /api/v1/admin/pai/keys accepts optional per-key restrictions:
{
"name": "Internal Analytics",
"ip_allowlist": ["10.0.0.0/8", "192.168.1.10"],
"domain_allowlist": ["internal.example.com", "*.svc.example.com"]
}
For backwards compatibility, mixed entries submitted in ip_allowlist are split into IP/CIDR and domain restrictions by the backend. Domain restrictions are matched against X-Forwarded-Host, Host, Origin, and Referer; use them as an additional guard with the global IP allowlist, not as a replacement for network controls.
Email delivery is best-effort and is not required before PAI key issuance (GA-only generation succeeds even when SMTP is unavailable; the companion token is shown on screen and, when SMTP is configured, also emailed). MFA verification is planned but not yet available.
PAI_IP_ALLOWLIST_MODE, added 2026-08-23)PAI_IP_ALLOWLIST_MODE decides how every PAI IP allowlist is applied. It is an environment
variable — it cannot be changed at runtime or from the admin UI.
| Value | Behaviour |
|---|---|
enforce (default) | A client IP outside the allowlist is rejected with 403. Unchanged historical behaviour. |
monitor | The identical check runs, every miss is recorded and logged at WARNING, and the request is allowed through. Use it to trial an allowlist against real traffic before it starts rejecting callers. |
Any unrecognised value resolves to enforce: a typo must never silently switch an access control off.
The mode applies to all three IP layers — the system-level PAI_IP_ALLOWLIST, a key's own
ip_allowlist, and the per-key list carried inside a session JWT. It covers the IP axis
only: domain allowlists, companion tokens, scope checks and rate limits are unaffected, so a request
that fails any of those is still rejected in monitor mode.
An empty allowlist remains a no-op in both modes — “not configured” has always meant
“not applied”, and monitor does not invent misses for a layer that was never
restricted. While monitor is active a configured allowlist is not a control, so the process-start
“PAI is unrestricted” warning still fires.
GET /api/v1/admin/pai/ip-allowlistGlobal-admin only. Read-only report of how the allowlist is currently applied, backing the IP Allowlist panel in the admin dashboard. Returns no key material.
{
"mode": "monitor",
"raw_mode": "monitor",
"effective": "monitoring",
"system_allowlist": ["10.0.0.0/8"],
"trusted_proxy_hops": 2,
"require_paired_token": false,
"key_mode": "db",
"caller_ip": "203.0.113.5",
"caller_allowed": false,
"observations": [
{
"observed_at": "2026-08-23T18:05:11.412000+00:00",
"layer": "system",
"client_ip": "203.0.113.5",
"allowlist": ["10.0.0.0/8"],
"path": "/internal/v1/cve/CVE-2024-0001",
"method": "GET"
}
],
"observation_total": 1,
"observation_retained": 1,
"observation_limit": 100
}
effective is not_configured when the allowlist is empty, otherwise
enforcing or monitoring. caller_allowed is null when
nothing is configured, since “allowed” would imply a check ran. layer is one of
system, key, or session_key.
Observations are an in-memory, per-worker ring buffer capped at
observation_limit — a rollout aid, not an audit log. A multi-worker deployment shows only
the worker that served the request, and a restart clears it. The WARNING log line is the durable
record.
DELETE /api/v1/admin/pai/ip-allowlist/observationsGlobal-admin only. Clears the retained observation buffer and returns
{"cleared": true, "dropped": <n>}. Application logs are unaffected. The action is written
to the admin audit log.
x-pai-token, added 2026-07-21)
Database-mode PAI keys (pai_key_mode=db) can carry an optional companion token — an
independent secret (phx_pait_...) minted alongside the api key by
POST /api/v1/admin/pai/keys (returned inline as pai_token in that response, and
also emailed to the requesting admin). Send it as a second header alongside the key:
x-pai-token: phx_pait_...The companion is required — a missing or mismatched value returns 401 — when either:
pai_require_paired_token is true (env PAI_REQUIRE_PAIRED_TOKEN, default false), or
This is a backward-compatible, per-key rollout: existing db-mode keys issued before this feature (no
stored companion hash) keep working with the api key alone unless pai_require_paired_token
is turned on globally. Env-mode PAI keys (pai_key_mode=env, the single shared
PAI_KEY) never require a companion token — there is no per-key record to hang a hash
off of.
Authorization: Bearer <jwt>, added 2026-07-21)
Once minted via POST /internal/v1/pai/session (below), a short-lived session JWT can be
presented as Authorization: Bearer <token> in place of the api-key + companion-token
pair on subsequent internal API calls, until it expires or the key's revocation epoch advances.
POST /internal/v1/pai/session
Exchanges a PAI api-key + companion-token pair for a short-lived session JWT, so subsequent calls can
send Authorization: Bearer <token> instead of resending the raw key pair on every
request.
Flag: pai_session_tokens_enabled (default false, env
PAI_SESSION_TOKENS_ENABLED). Returns 404 when off.
Auth: Same as any PAI call — the router-level require_pai_access
dependency runs first, so the IP allowlist, api key, and (if required for this key) companion token are
already validated before this handler runs. If that validation fails, the usual 401/403
responses apply and the request never reaches the session-minting logic below.
No request body. Send the pair as headers:
POST /internal/v1/pai/session X-API-Key: phx_pai_... x-pai-token: phx_pait_...
(Omit x-pai-token if the key doesn't require a companion — see Paired companion token above.)
{
"access_token": "<jwt>",
"token_type": "Bearer",
"expires_in": 900
}
expires_in is seconds until expiry, from pai_session_ttl_seconds (default
900, i.e. 15 minutes; env PAI_SESSION_TTL_SECONDS).
Present it on subsequent internal API calls in place of the key pair:
Authorization: Bearer <access_token>
The minted JWT embeds the same per-key IP/domain allowlist as the underlying key, so the session path is
never broader than what the raw api-key + companion pair could already reach. It is also invalidated
early if the key's revocation epoch advances (e.g. the key is revoked/rotated) —
verify_session_token checks the JWT's embedded epoch against the key's current epoch on
every call and rejects a stale one.
Bearer session token to this endpoint returns 401 ("Cannot mint a
session from an existing session token..."). You must always present the api-key + companion pair to
mint a session; this prevents a leaked short-lived token from self-refreshing indefinitely.401 Missing PAI key, 401 Invalid PAI key,
401 Missing PAI companion token, 401 Invalid PAI companion token,
403 Client IP not allowed) apply as usual if the pair fails validation before reaching
this endpoint.| Endpoint | Purpose |
|---|---|
POST /internal/v1/pai/session | Mint a short-lived session JWT (see Session Tokens above) |
GET /internal/v1/cve/{cve_id} | Full CVE details + raw NVD record |
GET /internal/v1/cves/{cve_id}/patch-window | Full, unredacted patch-window record (mirror of GET /api/v1/cves/{cve_id}/patch-window); gated by enable_patch_windows, 404 when off |
GET /internal/v1/cve/{cve_id}/github-pocs | Full (uncapped) GitHub PoC list + PS-HP-aligned popularity summary |
GET /internal/v1/phoenix-score/{cve_id} | Full PS-HP output with components |
GET /internal/v1/high-profile | Full high-profile list (no redaction) |
GET /internal/v1/enterprise-watchlist | Full watchlist entries |
POST /internal/v1/packages/intel | Internal package-version CVE + MPI malware intelligence |
POST /internal/v1/calculate-score | Custom PS-HP calculation |
POST /internal/v1/calculate-pss-score | Custom PS-OSS (PSS) calculation |
POST /internal/v1/calculate-phs-score | Custom PS-PHS calculation |
POST /internal/v1/calculate-pvs-score | Custom PS-PVS calculation |
POST /internal/v1/calculate-adqe-score | Custom ADQE calculation |
POST /internal/v1/calculate-eol-risk | Custom EOL SLR risk calculation |
GET /internal/v1/scoring-weights | PS-HP/PS-OSS weights |
GET /internal/v1/threat-actors | Threat actor intelligence |
GET /internal/v1/eol-intelligence | Full EOL intelligence |
GET /internal/v1/analysis/cwe | Full CWE analysis (mirror of /api/v1/analysis/cwe) |
GET /internal/v1/analysis/cwe-top25 | Full CWE Top 25 analysis |
GET /internal/v1/analysis/kev-cwe | Full KEV CWE analysis |
GET /internal/v1/analysis/owasp-top10 | Full OWASP Top 10 analysis |
GET /internal/v1/analysis/reference-threat-frequencies | Full, live reference threat-frequency distributions |
GET /internal/v1/analysis/cwe-intelligence | Full, unredacted merged CWE-intelligence table (mirror of /api/v1/analysis/cwe-intelligence); gated by enable_cwe_intelligence_endpoint, 404 when off |
GET /internal/v1/analysis/cwe/{cwe_id} | Full, unredacted single-CWE lookup (mirror of /api/v1/analysis/cwe/{cwe_id}); gated by enable_cwe_intelligence_endpoint and enable_cwe_catalog_intelligence, 404 when either is off |
POST /internal/v1/cwe/batch | Batch single-CWE lookup (1-500 ids; items keyed by id, same shape as GET /internal/v1/analysis/cwe/{cwe_id}); gated by enable_cwe_intelligence_endpoint and enable_cwe_catalog_intelligence, 404 when either is off |
GET /internal/v1/cves | Unrestricted CVE search — public filters plus sort, published_after/published_before, has_analysis; limit up to 1000; unredacted |
GET /internal/v1/cves/by-product | Unrestricted vendor/product CVE enumeration (limit up to 2000); response includes an additive cpe_uris field (sorted, deduplicated CPE 2.3 URIs from NVD) |
GET /internal/v1/cves/trending | Trending CVEs (news/social mention volume) |
GET /internal/v1/cves/kev | CISA KEV catalog, no tier field filtering (limit up to 10000) |
GET /internal/v1/cves/homepage-intel | Homepage intelligence bundle |
POST /internal/v1/cves/batch | Unrestricted batch CVE lookup (mirror of POST /api/v1/cves/batch); up to 5000 ids per call, unredacted; gated by enable_cve_batch_lookup |
GET /internal/v1/intel/entities | List the 8 queryable intelligence entity names (Unified Intelligence Query Surface mirror) |
GET /internal/v1/intel/{entity} | Search one entity, full fidelity, no data_shield shaping |
GET /internal/v1/intel/{entity}/{entity_id} | Get one entity by ID, full fidelity, no data_shield shaping |
POST /internal/v1/grants/trial, POST /internal/v1/grants/{grant_id}/convert, GET /internal/v1/grants/{grant_id} | Orange trial organization grants; gated by enable_pai_org_grants + enable_global_intel_license, 404 when off (see Organization grants) |
GET /internal/v1/cve/{cve_id} does NOT carry the public route's additive CVE-detail
blocks. It is a structurally different code path (InternalDataService.get_full_cve)
that returns {"cve": {...raw enriched dict...}, "nvd_raw": {...}} directly — it never
calls from_full_cve, _apply_purple_enrichment, the ATT&CK-technique lookup, or
the new CWE-catalogue lookup. Concretely, this PAI route does not gain the
cwe_intel block added 2026-09-06 to the public
GET /api/v1/cves/{cve_id} (see PUBLIC_API.html). The CVE's raw
cwes id list is already present, unredacted, inside the returned cve object; it is
simply never cross-referenced against cwe_catalog.json the way the public route's block does. A
PAI caller wanting the catalogue name/description/taxonomy for one of those CWE ids should call
GET /internal/v1/analysis/cwe/{cwe_id}
directly, per CWE id. Not an oversight: this route predates the CWE-catalogue feature and was already
fully unrestricted, so there is no tier gate here to bypass — the enrichment was simply never wired
into this call path.
The /internal/v1/cves* group (added 2026-07-08) mirrors the public
/api/v1/cves read surface, unredacted and with higher limits, gated only by
PAI auth (require_pai_access + pai_enabled).
POST /internal/v1/cves/batch (added 2026-08-20) mirrors the public
POST /api/v1/cves/batch (Purple R1): same id-normalisation (uppercase,
dedupe, sort, silently drop malformed ids — all-malformed is a 422, not
an empty 200) and the same explicit not_found[] list, but returns records
unredacted (no tier projection) and accepts up to a flat 5000 ids per
call instead of the public route's per-tier cap. Gated on the same
enable_cve_batch_lookup flag as the public route (default false,
returns 404 when off) so both transports move together.
not_found semantics. An id lands in not_found when the
corpus has no NVD record for it — determined by the returned record carrying no NVD-derived
content (no description, published, modified or
cvss). This matters because the underlying get_enriched_cve() never
signals a miss: it returns a populated-looking envelope (cve_id plus an
enrichment block) for any id, including one that does not exist. Treat
items as authoritative and not_found as "not in Blue's corpus", rather
than assuming every requested id comes back in items.
One known imprecision, stated so it is not mistaken for corruption: a CVE that genuinely exists in
NVD but has no description, no dates and no CVSS — a freshly reserved or rejected entry
— is reported in not_found. No fabricated record is ever returned in
items.
Mirrors the public Patch-Window Remediation surface
(Tasks 13–16). Gated only by enable_patch_windows (404, never 403, when
off) — there is no tier check on this plane: PAI is unshaped and first-party, so a
caller who clears require_pai_access sees the complete, unredacted record, identical to what a
Pro/Enterprise public caller sees after data_shield filtering (every patch_window_*
leaf is already classified PRO, and PRO is the record's full field set — filtering removes nothing at
this tier, it is simply never applied here).
GET /internal/v1/cve/{cve_id} — cve.patch_window is
attached by InternalDataService.get_full_cve whenever the flag is on and a
record exists for the CVE; otherwise the key is absent, never null.GET /internal/v1/cves/{cve_id}/patch-window — the dedicated per-CVE
mirror. Checks the flag first and returns 404 before any lookup; a CVE with no record is
also 404. Returns {"cve_id": ..., "patch_window": {...}}, the raw record from
patch_windows/by_cve.json, every leaf name verbatim — nothing renamed or
restructured relative to the public shape.POST /internal/v1/cves/batch — opt-in only, via the
include-token mechanism documented below: sending "include": ["patch_window"]
(only valid when enable_patch_windows is on, otherwise 400 as an unknown
token) attaches the record under items[i].intel.patch_window, every leaf name
verbatim. A request with no include never adds this block. The envelope
shape ({total, found, not_found, items}) is unchanged either way. (Earlier revisions
attached an unconditional top-level patch_window key here regardless of
include — removed 2026-09-26, D19 merge note in CHANGELOG.md.)
There is no PAI mirror of the public month-view endpoint
(GET /api/v1/patch-windows/{vendor}/{window_id}) — the per-vendor window data is already
fully available unshaped through the
patch_window bulk export dataset
below, the intended first-party bulk-consumption path.
POST /internal/v1/packages/intel
Returns the internal package-version intelligence payload for a package coordinate. This is the PAI twin of the premium public package intelligence endpoint.
Gated by enable_package_intel_api (default false). When the flag is off
the endpoint returns 404.
{
"ecosystem": "Maven",
"name": "com.h2database:h2",
"version": "1.3.176"
}
| Field | Notes |
|---|---|
ecosystem, name, version, purl | Normalised package identity. |
cves | Version-filtered OSV CVE/advisory records. |
vulnerabilities | Alias of cves. |
malware | MPI and compromised-package intelligence. |
malware.scan_id, malware.signals, malware.iocs, malware.raw | Internal-only MPI evidence for Phoenix-owned services. |
ps_oss_score | Derived package risk score. |
source_attribution | Data sources used in the response. |
Gated by ENABLE_ADV_FETCH (default false). All return 404 when the flag is off.
| Endpoint | Purpose |
|---|---|
GET /internal/v1/cves/{cve_id}/advisory-patches | Advisory patch suggestions for a CVE |
GET /internal/v1/advisories/{advisory_id} | CVEs + full advisory records for an advisory ID (422 on malformed ID) |
GET /internal/v1/advisories/patch-month/{patch_month} | Advisory entries for a patch month (YYYY-MM; 422 on bad format, 400 when > 500 records) |
Added 2026-07-10. PAI mirrors of the five public /api/v1/analysis/* endpoints (see
PUBLIC_API.html). These routes are not themselves gated by
ENABLE_CWE_OWASP_THREAT_TIERED_API — that flag only governs whether the public
routes shape their payload. The PAI mirrors always return the full, unshaped payload
(AccessMode.FULL), regardless of the flag's state, because PAI access is itself the Enterprise
gate (require_pai_access) and there is no end-user tier context on a PAI call.
| Endpoint | Purpose |
|---|---|
GET /internal/v1/analysis/cwe | Full CWE analysis (mirror of /api/v1/analysis/cwe) |
GET /internal/v1/analysis/cwe-top25 | Full CWE Top 25 analysis |
GET /internal/v1/analysis/kev-cwe | Full KEV CWE analysis |
GET /internal/v1/analysis/owasp-top10 | Full OWASP Top 10 analysis |
GET /internal/v1/analysis/reference-threat-frequencies | Full, live reference threat-frequency distributions — note: this route always reads the true, current reference_threat_frequencies.json, unlike the public route, whose underlying data source is itself flag-gated (legacy snapshot when the flag is off). |
No additional tier parameter or header is required beyond the PAI key itself. Missing/unavailable
underlying data returns 404 (this router's existing convention — see
GET /internal/v1/eol-intelligence), not the public routes' 503.
GET /internal/v1/analysis/cwe-intelligence is the PAI mirror of the public
merged
CWE-intelligence table (backend/app/routers/analysis.py). Same outer-join merge over
/cwe, /kev-cwe, /owasp-top10 and CWE_MASTER, keyed by bare
CWE id.
Unlike the CWE/OWASP/threat-frequency mirrors above (not flag-gated — always full, gated only by
require_pai_access), this route is flag-gated:
enable_cwe_intelligence_endpoint (default false, absent from
site_config.json), same flag as the public route. Returns 404 on this PAI plane
when off, same as the public route (both are new routes with no legacy unflagged behavior to preserve).
Always AccessMode.FULL, passed explicitly at this one call site rather than baked into the shaper:
raw nvd_count and raw kev_count are present on every row. This is the
only plane where either raw count is ever returned — the public route's own tier ceiling (Enterprise
→ PERCENT_ONLY) never reaches FULL; see
PUBLIC_API.html for why the
public route uses a dedicated resolver instead of the shared resolve_access_mode.
Same response envelope as the public route: schema_version: "cwe-intel-v1", per-source
source_provenance list ({"name", "generated_at"}, named to avoid colliding with each
row's own differently-shaped sources: List[str] field), cwes (dict keyed by bare
CWE id), total_unique_cwes (== len(cwes)), dropped_keys (count of raw
source keys that failed CWE-id normalisation and were therefore excluded from the merge —
0 in the steady state). Same caching contract:
ETag is a sha256 over the shaped (AccessMode.FULL) bytes — computed the same
way as the public route's ETag over its own lower-tier bytes, so the two planes never collide on the same
digest for different content — plus Cache-Control: private, max-age=0, must-revalidate and
If-None-Match → 304 support.
?scope= on this plane (added 2026-09-06)
GET /internal/v1/analysis/cwe-intelligence?scope=<observed|catalog> behaves
identically to the public route's ?scope= (see
PUBLIC_API.html), including the flag
interaction — symmetric across both planes, not PAI-specific:
enable_cwe_catalog_intelligence off: scope accepted but forced
to observed (never 422'd for a valid value); cwe_catalog/cwe_statistics
are not even loaded; schema_version reports the legacy "cwe-intel-v1".enable_cwe_catalog_intelligence on: scope=catalog takes effect
(308 → 989 rows, same measured counts as the public route); schema_version reports
"cwe-intel-v2"; the new catalogue fields and a FULL-tier statistics
block (every raw field) are attached wherever a corresponding entry exists.scope value is rejected with 422 regardless of the flag's state.
enable_cwe_intelligence_endpoint (the route-level 404 gate above) and
enable_cwe_catalog_intelligence (this ?scope=/schema gate) are two independent
flags — the route can be reachable (first flag on) while scope/the new fields stay dark
(second flag off).
GET /internal/v1/analysis/cwe/{cwe_id} is the PAI mirror of the public
Single CWE Lookup. Brand-new
route: gated on both enable_cwe_intelligence_endpoint (the surface-wide
gate the table mirror above checks first) and
enable_cwe_catalog_intelligence; either one off is a 404. It originally
checked only the second, so with the surface-wide flag off this mirror returned a full 200 body while
/internal/v1/analysis/cwe-intelligence 404'd on the same request — corrected in the
final whole-branch review, 2026-09-06.
{cwe_id} accepts 79/CWE-79/cwe-79 (normalised via
normalise_cwe_key); an id that fails to normalise, or normalises but has no row anywhere in the
full catalogue, is also 404. Always merges at scope="catalog" internally (any
of the 969 catalogued CWEs) — there is no ?scope= parameter on this route on either plane.
Always AccessMode.FULL, passed explicitly at this one call site: raw nvd_count/
kev_count and a fully-raw statistics block are present whenever a corresponding
entry exists. Returns the single shaped row plus its own id field (normalised
CWE-{n} form) — no table envelope. Same ETag/
Cache-Control: private, max-age=0, must-revalidate pattern, computed over this row's own shaped
bytes.
POST /internal/v1/cwe/batch (added 2026-10-01)
Batch counterpart of the single lookup above, so a consumer refreshing a set of stale or missing CWE ids does
not re-pull the whole catalogue (ASPMAIN-7795). Request body:
{"cwe_ids": ["79", "CWE-89", "cwe-200"]} — 1 to 500 entries (422 above 500,
when empty, when not a list of strings, or when an entry is longer than 64 characters).
Same gates and shaping as the single lookup: both enable_cwe_intelligence_endpoint
and enable_cwe_catalog_intelligence must be on (either off is a 404), standing
require_pai_access, ids normalised via normalise_cwe_key
(79/CWE-79/cwe-79/" 79 " are one id),
scope="catalog" merge, AccessMode.FULL. Ids are deduplicated after normalisation,
first-seen order preserved.
Response: {"total": <distinct ids>, "found": <int>, "not_found": [...], "items": {"<id>": {...}}}.
items is keyed by the bare numeric id and each value is byte-for-byte the object
GET /internal/v1/analysis/cwe/{cwe_id} returns for that id (including its id field).
An id that does not normalise, or has no catalogue row, is listed in not_found (the normalised id
when it normalised, else the trimmed input) and never fails the call. No ETag (the body is a batch,
not a single row).
Added 2026-07-25. PAI mirror of the public
Unified Intelligence Query Surface
(backend/app/routers/intel_query_pai.py), over the same resolver registry
(backend/app/services/intel_query/registry.py) as every other adapter for this feature.
Gated by the same flag, enable_intel_query_surface (default false) —
returns 404 when off, in addition to the standing require_pai_access
dependency on the whole /internal/v1/intel/* router.
| Endpoint | Purpose |
|---|---|
GET /internal/v1/intel/entities | List the 8 queryable entity names |
GET /internal/v1/intel/{entity} | Search one entity (q, vendor, product, purl, limit 1–500 default 100, offset) |
GET /internal/v1/intel/{entity}/{entity_id} | Get one entity by ID |
Final whole-branch review fix-round (2026-07-25): the limit ceiling was
lowered from 1000 to 500 to match Enterprise's real max_page_size (tier_limits.py)
— PAI access is itself the Enterprise gate, so a caller could previously ask for double what even an
Enterprise public caller's own tier cap allows. The resolver call is also now offloaded to a worker
thread (anyio.to_thread.run_sync), matching the public route's own fix (see
PUBLIC_API.html's "Known rollout gates" section) — some resolvers
(CveResolver in particular) run unindexed O(N) scans over in-memory data and must never
block the event loop.
Full fidelity — no data_shield pass — for a phx_pai_
(pai_internal) key. Unlike the public route, which
runs exactly one filter_response_for_tier pass keyed to the caller's resolved tier, this
router runs no field-level shaping for an internal PAI key. Every field a resolver emits
is returned as-is. provenance.redaction_profile is stamped "pai" (not a tier
name) so a deterministic consumer can tell which shaping pass — none — produced the response.
As on every other PAI endpoint, PAI's advantage is field-level fidelity, never internal business data:
LLM model names and per-call costs are absent because no resolver emits them, not because they were
redacted.
A phx_gintel_ (api_global_intel) BUNDLE key IS shaped, at the tier its
licence grants (fix C1, 2026-08-14). The "PAI access is itself the Enterprise gate" premise
above holds for the internal key family (pai_internal maps to ENTERPRISE); it
was falsified for this router when Plan B admitted a customer key sold at two levels. A
global_intel:pro key was receiving the ENTERPRISE-classified
deterministic.signals / chains that the identical query at Pro on
/api/v1/intel/* strips. A bundle key's response — including the 403 body
— now goes through exactly one filter_response_for_tier pass at
enterprise / pro, floored at registered when the key carries no
valid stamp, and provenance.redaction_profile names that tier instead of "pai".
Same licence, same fields, whichever door. Its limit is likewise clamped to its own tier's
max_page_size rather than the Enterprise ceiling of 500.
Bundle-key admission is flag-gated (same fix). With
enable_global_intel_license off, the accepted-scope tuple collapses to
PAI_DEFAULT_SCOPES (pai_internal only) and a phx_gintel_ key gets
401 here too — so the documented rollback kill switch genuinely closes this door
instead of leaving already-issued bundle keys admitted on the one plane where the tier demotion had no
effect. An unreadable flag store means no admission.
allow_clean (malware verdict access gate) — PAI does NOT bypass it.
Read this before "fixing" it back to a full bypass.
Whether a caller can resolve an unconfirmed/under-review (SUSPECT-band) malware package at
all is a completely separate mechanism from the field-shaping described above (see
PUBLIC_API.html's malware verdict access gate section). This router never
passes allow_clean=True for the malware entity — it relies entirely on
resolve_entity/search_entity's own default (False), identical to
how an anonymous/non-admin caller is treated on the public route. A PAI caller sees confirmed-malicious
verdicts exactly the way a non-admin public caller does; clearing the PAI dependency does
not, by itself, grant visibility into scanned-but-not-confirmed packages.
This was an explicit, human-reviewed security decision made during implementation, reversing an earlier
draft that forced allow_clean=True unconditionally for PAI on the theory that
backend/app/routers/mpi_artifacts.py (a PAI-only, no-admin-stacking read router with no
verdict gate at all) was sufficient precedent that PAI-alone is this platform's ceiling of trust for MPI
data. That reasoning was rejected: loosening who can see unconfirmed/under-review SUSPECT-band malware
verdicts is an authorization-policy change requiring explicit approval before weakening any access gate
(per AGENTS.md Hard Rules), and backend/app/routers/malware_intel_internal.py
stacks an extra require_global_admin check on top of require_pai_access
specifically for other sensitive malware operations (force-rescan, raw repo_analysis blob)
— direct evidence that this codebase's convention does not treat PAI alone as
admin-equivalent for sensitive malware data.
If a future task needs PAI to see SUSPECT-band packages, it requires a distinct, higher-trust PAI signal
(e.g. a specific key scope or attribute) — not simply restoring the unconditional
allow_clean=True this router originally shipped with.
?include=advisory (Plan D, added 2026-07-26)
GET /internal/v1/intel/malware/{entity_id}?include=advisory (and the equivalent on
GET /internal/v1/intel/malware?include=advisory search) populates the envelope's
advisory block with MPI LLM reasoning —
{"analyst": {"verdict", "confidence_class", "reasoning"}, "judge": {"verdict", "reasoning"},
"reproducible": false}, built by the shared build_reasoning_view()
(app/services/mpi_reasoning.py). See
PUBLIC_API.html's "Advisory block opt-in" section for the full field-by-field
detail and the confidence-class vocabulary it uses.
PAI honors this unconditionally whenever the caller passes it — no additional tier
check, unlike the public REST mirror, which additionally requires the caller's resolved tier to be
enterprise. This is consistent with this route's "full fidelity, no
data_shield pass" behavior for an internal PAI key described above: PAI access is already this
platform's highest trust boundary, so there is nothing narrower to check.
For a phx_gintel_ bundle key the gate is still not applied, but the shield is (fix
C1, 2026-08-14). The call still returns 200 and still needs no bundle —
advisory is not an expanded token and never triggers the licence gate — but
advisory is ENTERPRISE in data_shield, so it survives only for a
global_intel:enterprise key. Below that it comes back null, exactly as it does
on /api/v1/intel/* at the same tier.
This does not widen the allow_clean verdict-visibility gate described
immediately above. include_advisory is only ever consulted by the malware resolver
after an allow_clean-gated row load has already succeeded — a PAI caller
requesting ?include=advisory for an unconfirmed/under-review (SUSPECT-band) package still gets
the same 404 any other non-privileged caller gets.
MalwareResolver.get() additionally nulls advisory —
independent of the allow_clean-gated row load above — for a D-14-eligible-but-non-terminal-
verdict row (freshest scan still SUSPECT/AWAITING_REVIEW/LLM_ANALYSIS/
INCONCLUSIVE) unless the caller is exempt; the row itself still resolves (200), only
advisory comes back null. On REST this applies to every caller except a global
admin. This router passes a distinct bypass_advisory_verdict_gate=True (never
allow_clean=True, which would incorrectly reopen the gate described above) alongside
include_advisory=True, so PAI keeps the fully unconditional advisory behavior documented above
regardless of verdict state. A first implementation of the terminal-verdict check keyed its exemption off
allow_clean alone, which silently narrowed PAI too since this router never sets
allow_clean; this was caught and fixed before it shipped as a behavior gap.
Requires both enable_intel_query_surface and
enable_mpi_reasoning_enterprise to be on; advisory stays null if
either is off.
?include= (Plan B, added 2026-08-13)
GET /internal/v1/intel/library_version/{purl}?include=vulnerabilities,malware,exploitation,campaign,licensing
returns the library, its vulnerabilities with per-CVE detail, its malware intelligence, and its
exploitation / campaign / licence vectors in one call, under a new top-level
expanded envelope block. GET /internal/v1/intel/library/{purl} (unversioned)
accepts the same tokens and additionally returns a cross-version rollup
(any_version_compromised, compromised_versions[], safe_versions[],
worst verdict).
PAI first, then everywhere (Plan D, 2026-08-14). PAI shipped first because it
skipped data_shield entirely and stamped redaction_profile="pai", so the
composition shape could be proven before being locked behind a tier matrix. Plan D added the same
?include= surface to public REST (/api/v1/intel/*), the MCP
intel_get tool and phoenix-cli, all four sharing the registry's single token
gate. See PUBLIC_API for the tier matrix — which since fix C1 applies
to a phx_gintel_ caller on this router too, and is not applied only
to an internal phx_pai_ key.
SCHEMA_VERSION moves 1.0.0 → 1.1.0. The bump is
MINOR because expanded is opt-in: a request with no include returns a
byte-identical response — the expanded key is absent, not
null, and provenance keeps exactly its four original keys.
Authentication. Both phx_pai_ (pai_internal) and
phx_gintel_ (api_global_intel) keys authenticate on this router,
through the same validate_pai_request path: same system IP allowlist, same
per-key IP/domain allowlists, same companion-token requirement, same rate limiter. A Global Intel key is a
licence, not a way around PAI's network controls.
/internal/v1/intel/* only.
phx_gintel_ is a customer key — issued against an org's
global_intel product entitlement — while /internal/v1/* as a whole is the
privileged internal plane (MPI forensic artifact reads via mpi_artifacts.py, which is PAI-only
with no admin stacking at all; org/team tenancy administration; research triggers). Admitting the scope
globally in validate_pai_request would hand all of that to any Global Intel customer. So
PAI_DEFAULT_SCOPES stays (PAI_INTERNAL,) — unchanged for every other PAI
endpoint — and this router opts in via require_pai_or_global_intel /
PAI_INTEL_QUERY_SCOPES. A phx_gintel_ key presented to any other
/internal/v1/* endpoint gets 401, exactly as before.
include token | Global Intel Pro | Global Intel Enterprise |
|---|---|---|
vulnerabilities | yes, capped 100 | yes, capped 500 |
licensing |
spdx_id, category, license_risk_score, license_policy | same |
exploitation | exploitation_tier + full flags/block_counts + pss_score (PS-OSS package score, int 0-100, null when the package has no score) | + tte.{tte_hours,speed_pressure,ecosystem_cohort} |
malware | verdict + malware_signals + mitre | + chains + iocs + phx_neural_score |
campaign | compromised + per-campaign name/date/severity | + campaign_threat_actor, tags, repeat_offender (+ reason), incident_count_365d, days_since_last_incident, decay_phase |
patch_window (added 2026-09-23) | full, unredacted record(s) — PAI is unshaped | same as Pro |
advisory | unchanged — see the section above | unchanged — see the section above |
include accepts CSV (?include=a,b) or repetition
(?include=a&include=b). include=all expands server-side to every token the
caller's tier permits and never errors on an over-privileged token. Token validation and the token→tier
matrix live in backend/app/services/intel_query/registry.py — one gate for four
transports, not one per adapter. patch_window is additionally excluded from the valid-token set
(400, not merely licence-refused) whenever enable_patch_windows is off.
patch_window is new to COMPOSABLE_ENTITIES via cve, not only
library/library_version. cve was added solely to carry this
one token (2026-09-23) — no other token composes on a cve entity. Licence resolution is
entity-agnostic, so GET /internal/v1/intel/cve/{cve_id}?include=patch_window still needs the
Global Intel bundle, exactly like a library/library_version composition —
being pai_internal-authenticated does not itself bypass the bundle requirement for this token.
The output key differs by entity, deliberately. On the cve entity the composed
block is the flat key expanded.patch_window (one record, same shape as the
Patch-Window Remediation mirror above). On
library/library_version it is expanded.patch_window_by_cve — a
CVE-id-keyed map, capped at 200 entries. patch_window_truncated is a sibling key inside that
SAME map (not a separate field), is always present — true once the cap is hit,
false otherwise, never absent — and MUST be skipped when iterating the map's keys as CVE
ids. See the renames table below for why the map shape could not reuse the bare
patch_window name.
Several envelope field names are deliberately not the source's own.
data_shield's DEFAULT_FIELD_CLASSIFICATIONS is a flat map keyed on bare field
names and filter_response recurses by bare key, so a nested key here shares one classification
with every dataset on the platform. Two distinct hazards force the renames, and both are live.
| Envelope name | Source name | Hazard |
|---|---|---|
malware.malware_signals | signals |
signals is ENTERPRISE for the MPI dossier; adopting it would ship a Pro tier that pays
for signals it cannot read, and relaxing it would leak dossier signals to Pro everywhere. |
malware.malware_signals[].malware_signal_id | signal_id |
Same. Without this one, a Pro caller receives [{"category", "severity"}] —
signal rows with no identity, which is worse than withholding the array because it looks like a
feature. |
malware.mitre.mitre_technique_ids | technique_ids |
Same. |
licensing (block + include token) | license |
license is currently unclassified, and under
data_shield_fail_closed unclassified means dropped.
description_sources[].license ("CC-BY-4.0") is dropped platform-wide today on
the cve dataset; classifying the bare name would newly expose it
there. |
licensing.license_policy / licensing.license_risk_score |
policy / risk_score |
Same class of hazard; renamed pre-emptively because the names are generic. |
exploitation.block_counts | counts |
enrichment.circl_history.counts on the cve dataset — same inverse
hazard as license. |
exploitation.tte.tte_hours | hours |
Generic; renamed pre-emptively. |
campaign.campaigns[].campaign_threat_actor | threat_actor |
projections.from_full_cve emits threat_actor on the cve
dataset. |
provenance.intel_license | license |
As licensing above. |
library/library_version's patch_window_by_cve |
patch_window (as a CVE-id-keyed map) |
A third, distinct hazard, found in review (2026-09-24): the bare name patch_window is
also used, unchanged, as a flat single-record object with real classified leaf names on six
other already-shipped surfaces (CVE detail, batch, the PAI mirror above, MCP, and the cve
entity's own expanded.patch_window). Registering patch_window as a
DYNAMIC_KEYED_MAP_FIELDS entry would have silently exempted every one of those flat objects
from field-name classification too. The ?include=patch_window token name is unchanged; only
the library/library_version output key was renamed. |
The second hazard is the less obvious one and is worth stating plainly: adding a
classification is not free either. Under fail-closed, an unclassified name is dropped for every tier, so
classifying a name that another data_shield-filtered producer also emits is a live behaviour
change to that unrelated, already-shipped endpoint.
tests/test_data_shield_intel_coverage.py (D-T2c) checks this mechanically against every module
reachable from a filter_response_for_tier call site. None of these renames is cosmetic.
| Condition | Behaviour |
|---|---|
No include | Byte-identical to the pre-Plan-B response. |
| Unknown token | 400, detail naming the unknown token(s) and listing every valid token. |
enable_intel_query_surface off | 404 — checked before include parsing and before licence resolution, so neither a 400 nor a 403 can leak the feature's existence. |
enable_global_intel_license off, or no bundle on the key | 403 {"detail": "Global Intel license required for expanded intelligence"} — and the base envelope is still in the response body. Losing the bundle narrows the answer; it does not delete it. |
Daily intel_expanded quota exhausted | 429 {"detail": "Usage limit exceeded"} (Plan C, added 2026-08-14). Checked after the licence gate and before composition runs, so a rate-limited call costs a dict lookup rather than a full fan-out. |
Metering and quota (Plan C, added 2026-08-14). Expanded composition is metered separately from a plain lookup because it is materially more expensive — it fans out to MPI, OSV/library-intel and up to 500 per-CVE exploitation lookups.
| Counter | Increments on |
|---|---|
intel_query_lookups | every billable call to /api/v1/intel/* or /internal/v1/intel/*, composed or not |
intel_expanded_lookups | additionally, every call whose include names at least one licence-gated token on a transport that can serve it |
?include=advisory alone is not an expanded call: it needs no bundle,
does not increment intel_expanded_lookups, and is not subject to the
intel_expanded quota. Only the tokens in EXPANDED_TOKENS (plus
all) trigger either — the classifier consults
registry.wants_expanded(), the same gate the licence check uses, so the two cannot
drift.
All three composing transports are billed identically (Plan D, 2026-08-14; corrected
2026-08-14 by fix I2). The rule is "a transport bills intel_expanded_lookups
exactly when it can actually serve the composition". For the two HTTP transports that is
usage_metering.COMPOSING_PATH_PREFIXES =
("/api/v1/intel/", "/internal/v1/intel/"). MCP is the third: it has no
billable path of its own (classify_request decides /api/v1/mcp in an earlier
branch), so IntelToolHandler records the same two counters itself, from the same
usage_metering.intel_usage_keys() definition the path classifier uses. Until that fix MCP
enforced the intel_expanded quota against a counter only REST and PAI ever wrote
— composition served, billing zero, the 2,000/day Pro cap unreachable through that door. An
expanded call now costs the same fan-out and the same counter through any of the
three.
MCP bills intel_query_lookups on intel_get/intel_search and
adds intel_expanded_lookups only when a licence was actually resolved and the composition
served; api_calls_total/api_calls_mcp continue to come from
usage_tracking_middleware, which does not skip /api/v1/mcp, so nothing is
double-counted.
Metering has a hard prerequisite, and the surface now refuses without it. Both
counters are columns added by db/migrations/123-global-intel-usage-counters.sql. With 123
unapplied the read path (SELECT *) silently returns 0 and the write path fails wholesale
into a swallowed exception — the whole surface unmetered and unquota'd, with only a
WARNING. app/services/usage_schema.py now checks the columns directly: logged
loudly at startup, and every licence-gated expanded call returns 503 (naming the
migration) rather than serving the platform's most expensive read unbilled. Deploy order is 122 →
123 → code.
One consequence on the public route, and the shape of it matters: an
anonymous caller sending a licence-gated include is refused by the
route, not by _enforce_anonymous_limits, so the 403 still
carries the base envelope. Plan C had put intel_expanded in
ANONYMOUS_ENFORCEABLE_FEATURES so the middleware would refuse it before any resolver
work; once the public router actually composed, that became a regression — the middleware has
no envelope, licence or tier context, so all it could return was a bodiless
{"detail": …}, deleting data the caller was entitled to.
anonymous_feature_key now falls the anonymous enforcement back to search,
so the call is still metered and still 429-able while the route owns the refusal.
?include=advisory alone is unaffected: it needs no bundle and was never an expanded
call.
Daily intel_expanded caps: Pro 2,000,
Enterprise unlimited; free/registered are 0,
though an unlicensed caller is refused by the 403 licence gate before the quota is ever consulted.
Source of record is the tier_features DB table
(db/init/07-user-tiers.sql), mirrored by the fallback map in
backend/app/services/tier_limits.py; tests/services/test_tier_limits.py
pins the two together.
/entities and /vocabularies are static vocabulary listings and are
not metered. /internal/* is skipped entirely by
usage_tracking_middleware, so this router records its own usage — there is no
second writer and nothing to double-count. If the metering backend is unreachable the quota check
fails open (logged at WARNING): the licence gate is the authorization control, and
a storage blip must not deny a paying customer access they hold.
allow_clean is NOT widened by a Global Intel licence. This is the security
invariant of the whole feature. Composition reuses MalwareResolver in-process with
allow_clean at its default False, so a phx_gintel_
Enterprise key gets no expanded.malware block for a
SUSPECT-band (scanned-but-not-confirmed) package — not a redacted one, not an empty one. A product
entitlement is not the separately-decided authorization to see pre-confirmation triage state; see the
allow_clean section above for that decision's full record.
expanded.malware.verdict is never inferred, and absence of evidence is never
reported as cleanliness. A verdict may only come from a source the caller is authorized to see:
MalwareResolver (allow_clean-gated) for any verdict including
CLEAN; or the library_intelligence.json D-14 compromise flag for
MALICIOUS only, and only on the unversioned library entity where
resolve_purl_to_versioned can legitimately miss a purl_base that
compromised_package_intel knows about (that flag is already public to non-admins and cannot
encode SUSPECT-band state). Anything else — clean, gated, never scanned, or SUSPECT-band, which are
indistinguishable from outside the resolver — omits the block with
block_errors.malware = "malware_intelligence_unavailable". A caller who cannot be told "this
is clean" is told "we cannot answer", never "this is clean".
Provenance. When composition runs, provenance gains three keys, and a fourth
when there is something to disclose:
block_sources — per-block source attribution.
vulnerabilities is attributed to library-intel, not
osv: its rows come from the precomputed artifact, whereas
deterministic.vulnerability.cve_ids in the same envelope is a live OSV read. Only
malware is ever live.block_freshness — "live" for a block genuinely served from a live MPI
read; the library_intelligence.json artifact's _metadata.generated_at for
anything precomputed, so a stale enrichment lane is visible instead of presented as live.block_errors — {block: reason} for any omitted block. A failed or
timed-out block is omitted with a reason, never returned empty (an empty block reads as
"we looked and found nothing" when the truth is "we could not look") and never fails the whole call.block_notes — {block: caveat}, present only when a block that
did land carries one. Today the sole entry is vulnerabilities, disclosing that it
is a precomputed snapshot which may diverge from deterministic.vulnerability.cve_ids. When
the two disagree, treat deterministic as the live answer and
expanded.vulnerabilities as a snapshot of the age given in
block_freshness.
provenance.intel_license is stamped
{"product": "global_intel", "tier": "pro|enterprise", "source": "api_key:global_intel_tier"}
on a licensed response.
Truncation is never silent. expanded.vulnerabilities carries
total (the real count), returned, and truncated. Selection under the
per-tier cap is by severity then EPSS, so a cap never drops the most exploitable rows; presentation order is
then imposed by sort_deterministic() (determinism rule D5).
Array ordering (D5), precisely. sort_deterministic() orders a list of
objects by the first present key in envelope.STABLE_ID_KEYS (id,
signal_id, malware_signal_id, cve_id, cwe_id,
purl, cpe_uri, technique_id); lists of scalars sort naturally.
expanded.campaign.campaigns[] carries none of those keys and is therefore
ordered by its producer instead — ascending by (date, name) — which is stable,
tier-independent and reproducible, but is not the stable-ID rule the sentence above describes.
It was unordered entirely until 2026-08-14. advisory is exempt from D5 altogether: LLM
output is not reproducible and ordering it would imply that it is.
Budget. Each block has its own timeout,
INTEL_EXPANDED_BLOCK_TIMEOUT_MS (default 3000, capped at 60000). The whole resolve
is offloaded via anyio.to_thread.run_sync so composition never runs on the event loop.
Added 2026-07-26. PAI mirror of the public
Tenant-Scoped Research
Triggers (backend/app/routers/research_pai.py), over the same job store and
dispatcher as the REST surface. Gated by the same flag, enable_intel_research_triggers
(default false); returns 404 when off, in addition to the standing
require_pai_access dependency on the whole /internal/v1/research/* router.
POST /internal/v1/research — trigger research. Body: {"kind": "...", "target": "..."}. Returns 202.GET /internal/v1/research/{job_id} — get one job (org-scoped; another org's job reads as 404).GET /internal/v1/research — list this organisation's jobs (state, limit 1–200 default 50, offset).
org_id comes from the PAI key (PAIKeyInfo.org_id), never from a
request field — a PAI caller may not name an arbitrary org, or PAI becomes a
cross-tenant write primitive. A PAI key with no resolvable org_id gets 403.
No separate Pro/Enterprise tier check. A PAI key has no subscription tier to
resolve; quota is pinned at the Enterprise daily ceiling (250/day) rather than
gated on a lookup that does not exist for this transport — the same precedent as this
file's own Unified Intelligence Query Surface mirror capping its limit at
Enterprise's max_page_size instead of inventing a separate PAI-only ceiling.
This 250/day ceiling shares the SAME per-org daily counter as REST — it is
not an independent PAI-only allowance. create_job()'s quota check counts every row in
intel_research_jobs for the org created that UTC day, regardless of which transport
created it, so a Pro org's own REST quota is 25/day but a PAI key acting on that same org can
create up to 250/day against the identical counter — heavy PAI usage can exhaust the shared
counter and 429 the org's own subsequent REST calls for the rest of the day. This is a
deliberate, already-tracked rollout gate, not an oversight.
job_state, not state, in the response — the same
rename, and the same field-classification-collision rationale, as the REST mirror (see
PUBLIC_API.html). This
router duplicates a small, route-local _to_response() rather than importing
research.py's, matching this codebase's established PAI-adapter-router convention.
Feature doc: 2026-07-25-tenant-scoped-research-triggers.md.
Flag: enable_pai_bulk_export (default false). Every endpoint in this section returns 404 when the flag is off.
Serves pre-built, per-tier NDJSON snapshots of Phoenix datasets (backend/app/routers/bulk_export_pai.py).
Artifacts are produced out-of-band by pipeline/build_bulk_exports.py and served as-is — the
backend does not filter response data at request time, because the producer already wrote
base and premium as physically separate files (data_shield cannot scale to
millions-of-record exports).
Authentication. This surface admits PAI_INTERNAL keys only
(PAI_EXPORT_SCOPES in pai_auth.py) — narrower than the rest of PAI. A
phx_gintel_ (Global Intel bundle) key is rejected here even though it is admitted on
/internal/v1/intel/*: global_intel is a paid customer entitlement, and this surface's
licensing rationale holds only because every consumer is a first-party Phoenix platform.
global_intel may lift the tier of an already-admitted key
(resolve_export_tier); it never grants admission to this router.
Phase 3 delta is partial, not universal. Most datasets still carry
"delta": {"supported": false} in the manifest and must be re-pulled in full on every rebuild.
As of 2026-10-01 three datasets, patch_window, cve_intel and malware_packages, opt in
(DatasetSpec.delta_enabled=True, delta_retention=30) and additionally expose
GET /internal/v1/export/{dataset}/delta (below). A dataset with "supported": false
has no delta route at all — requesting one returns the same 404 as an unknown dataset.
GET /internal/v1/export/manifestLists every dataset this export generation produced, tier-shaped for the caller.
Free — it does not consume the bulk_export usage counter or the per-key
concurrency slot described below, so discovery never costs a consumer anything (a metered
discovery call would push consumers toward guessing dataset names instead of listing them).
{
"schema_version": "1.0",
"export_version": "v1",
"generated_at": "2026-08-19T02:14:07Z",
"tier": "base",
"datasets": [
{
"dataset": "kev",
"tier_required": "base",
"accessible": true,
"delta": {"supported": false},
"source_freshness": {},
"snapshot": {
"watermark": "2026-08-19T02:00:00Z",
"url": "/internal/v1/export/kev/snapshot",
"records": 1284,
"bytes": 483920,
"sha256": "deadbeef...",
"sha256_uncompressed": "c0ffee...",
"cursor": "eyJkIjoia2V2IiwidCI6ImJhc2UiLCJzIjo0MTJ9",
"snapshot_id": "kev-base-000412",
"source_freshness": {}
}
},
{
"dataset": "malware_packages",
"tier_required": "premium",
"accessible": false,
"delta": {"supported": true, "max_window_seqs": 30},
"source_freshness": {}
},
{
"dataset": "patch_window",
"tier_required": "premium",
"accessible": true,
"delta": {
"supported": true,
"max_window_seqs": 30,
"latest_cursor": "eyJkIjoicGF0Y2hfd2luZG93IiwidCI6InByZW1pdW0iLCJzIjo0MX0",
"retention_cursor": "eyJkIjoicGF0Y2hfd2luZG93IiwidCI6InByZW1pdW0iLCJzIjoxMX0"
},
"source_freshness": {
"patch_window_generated_at": "2026-09-26T04:00:00Z",
"patch_window_sources": [{"vendor": "microsoft", "fetched_at": "2026-09-26T04:00:00Z"}]
},
"snapshot": {
"watermark": "2026-09-26T04:00:00Z",
"url": "/internal/v1/export/patch_window/snapshot",
"records": 64000,
"bytes": 18400000,
"sha256": "deadbeef...",
"sha256_uncompressed": "c0ffee...",
"cursor": "eyJkIjoicGF0Y2hfd2luZG93IiwidCI6InByZW1pdW0iLCJzIjo0MX0",
"snapshot_id": "patch_window-premium-000041",
"source_freshness": {
"patch_window_generated_at": "2026-09-26T04:00:00Z",
"patch_window_sources": [{"vendor": "microsoft", "fetched_at": "2026-09-26T04:00:00Z"}]
}
}
}
]
}
The build registers exactly fifteen datasets — cve_core,
kev, epss, ransomware, exploitation_evidence,
library_intel, cwe, cve_taxonomy, patch_window,
cve_intel, product_health, product_vendor, phoenix_scores,
malware_packages, library_license (pipeline/bulk_export/datasets.py).
malware_packages (added 2026-10-01, ASPMAIN-7802) is shown above as a base caller
sees it. phoenix_scores, malware_packages and library_license declare
no required_flag: they are gated by enable_pai_bulk_export and the
premium tier only.
patch_window carries its own required_flag gate, fixed
2026-09-24 — DatasetSpec.required_flag="enable_patch_windows" is checked in
bulk_export_pai.py in addition to the whole-surface enable_pai_bulk_export flag
and the premium tier requirement: when enable_patch_windows is off,
patch_window is omitted from this manifest entirely (not listed, not merely
accessible: false) and its snapshot route 404s, regardless of caller tier. Originally
shipped (Task 19, 2026-09-23) with no per-dataset flag hook at all — a real
pai_internal+premium caller could list and download real patch-window data via
this surface even while the API-facing flag was off, silently defeating its own kill switch the moment
enable_pai_bulk_export was turned on. required_flag is generic (not hardcoded to
patch_window), so any future dataset whose producer can run ahead of its own API-facing flag
can reuse the same mechanism; every dataset registered before patch_window has
required_flag=None, a true no-op. cve_intel (2026-10-01) reuses it with
required_flag="enable_pai_cve_intel_dataset", resolved via feature_status.resolve_flag
(only enable_patch_windows is read straight off Settings). product_health
(2026-10-01, ASPMAIN-7800) reuses it with required_flag="enable_pai_product_health" — the
existing flag of POST /internal/v1/products/health, not a new one. product_vendor (2026-10-01, ASPMAIN-7801) reuses it with
required_flag="enable_pai_product_resolve_vendor" — the existing flag of
POST /internal/v1/products/resolve-vendor. malware_packages above is shown to illustrate
the shape of an inaccessible entry: a premium dataset stays listed
for a base caller with accessible: false and no snapshot block. Note
source_freshness is present-but-empty rather than absent:
shape_manifest_for_tier emits that key for every dataset entry,
accessible or not.
tier — the caller's resolved export tier (base or premium; see
resolve_export_tier). Starts at base and is max-at-resolution — never
lowers, never persists a derived row. On the live request path only a
global_intel stamp lifts a caller to premium.
resolve_export_tier also reads a per-key premium export-tier stamp and the org's
mpi entitlement, but neither is reachable today: PAIKeyInfo has no
export_tier field and nothing sets one, and bulk_export_pai._caller_and_tier calls
the resolver without an entitlement_lookup. Both branches are unit-tested and dormant until a
caller supplies them — do not document them as live inputs.accessible — whether the CALLER's tier satisfies tier_required for that
dataset. A premium-only dataset stays listed for a base caller with
accessible: false rather than being hidden — hiding it would make the 403 on a snapshot
request (below) incomprehensible, and the dataset's existence is a published product, not a secret.snapshot — present only when accessible is true;
snapshot.url is always /internal/v1/export/{dataset}/snapshot.
snapshot.cursor — the same opaque, resumable position as the
X-Phoenix-Export-Cursor response header on a direct snapshot request for this
dataset+tier (kept in sync). This is what a consumer persists;
snapshot.watermark is informational only (R2 — never a wall-clock range key).snapshot.snapshot_id — human-readable label for this build
({dataset}-{tier}-{seq:06d}), matching X-Phoenix-Export-Snapshot-Id.
Convenient for logs/support, not a substitute for cursor.snapshot.source_freshness (added 2026-10-01, additive) — the source
freshness recorded for THIS tier's artifact, in that artifact's own latest.json,
at the moment its source was read for the build. It is exact for the bytes this caller
downloads. It is {} for artifacts built before the field existed, and for
datasets whose serialiser does not populate freshness.source_freshness — per-source freshness metadata for this dataset
(DatasetSpec.source_freshness()), most relevant for a dataset that merges multiple
upstream artifacts of differing age (e.g. exploitation_evidence).
Three datasets populate it today (plus cve_intel, which returns one
<source>_generated_at file mtime per source plus news_window_date, and
product_health, which returns product_health_index_generated_at /
product_health_index_nvd_generated_at from the index plus
ps_hp_internal_generated_at / ps_hp_legacy_generated_at PS-HP file mtimes, and
product_vendor, which returns the same two product_health_index_* values, and
phoenix_scores, malware_packages and library_license, which return the same
{"generated_at": ...} as library_intel because they read the same file).
library_intel returns
{"generated_at": <library_intelligence.json's own _metadata.generated_at>}.
cwe returns {"generated_at", "catalog_version", "catalog_date"}
— MITRE's own version and release date alongside Phoenix's build time, so a consumer can
tell a rebuild that merely re-ran from one that picked up a NEW upstream catalogue, which
generated_at alone cannot distinguish. cve_taxonomy returns the same three
catalogue values, prefixed cwe_catalog_. Since 2026-10-01 this dataset-level value never over-claims. If the
tiers were built from different source captures (for example one tier failed and kept an older
artifact, or a concurrent build advanced one tier), it is the OLDEST tier's capture. For artifacts
built before 2026-10-01 that carry no per-artifact provenance, this is best-effort until each
dataset+tier has been rebuilt once. Use snapshot.source_freshness for the exact
per-artifact value. Every other dataset still returns
{} — no serialiser has wired one yet — since the field exists so each
producer can populate it independently without another manifest-shape change.delta — whether GET /internal/v1/export/{dataset}/delta (below) exists for
this dataset. {"supported": false} for every dataset except patch_window, cve_intel and malware_packages today.
When supported: max_window_seqs (the dataset's delta_retention — the widest
since..until span one request may cover, and how many builds of delta files are
kept on disk); latest_cursor — present only when accessible is also
true (it lives beside snapshot, one per tier) — the same cursor as the
snapshot's X-Phoenix-Export-Cursor, usable as until; retention_cursor
— the oldest since this dataset+tier can still serve a delta from before the caller must
re-baseline from the snapshot (409 delta_cursor_expired otherwise).503 (with Retry-After: 3600) when this export generation has not finished
building yet (manifest.json absent) — never a bare empty 200, which would read
as "no datasets exist."cwe dataset (added 2026-09-07)tier_required: base, 969 records, keyed by the bare numeric
CWE id ("79", never "CWE-79"). Built by
pipeline/bulk_export/serialisers/cwe.py, which joins two artifacts:
| Source | Produced by | Rows |
|---|---|---|
web/data/cwe_catalog.json | pipeline/rebuild_cwe_catalog.py (MITRE catalogue 4.20, dated 2026-04-30) | 969 weaknesses |
web/data/cwe_statistics.json | pipeline/rebuild_cwe_statistics.py | 799 rows |
The catalogue row is the record; the statistics row is attached as a nested
statistics block.
Who can call this at all. Bulk export is a PAI-plane surface with no
public-plane equivalent. require_pai_export
(backend/app/services/pai_auth.py:587) admits
PAI_EXPORT_SCOPES = (APIKeyScope.PAI_INTERNAL,) and nothing else, so every
caller needs a pai_internal key regardless of account tier — a Free,
Registered, Pro or Enterprise account with only a public API key gets no access to
cwe here and must use GET /api/v1/analysis/cwe-intelligence on the
public plane instead. A phx_gintel_ Global Intel key is explicitly
rejected for admission: the no-redistribution rationale holds only while every
consumer is first-party. Provisioning a pai_internal key is an operator action in the
admin dashboard, not a self-service upgrade.
Which half of each record you get. Admission and tier are two separate
decisions. resolve_export_tier
(backend/app/services/tier_resolver.py:339) runs only after a key is already
admitted, starts at base, and is max-at-resolution — it never lowers and never
writes a derived row.
| Caller | Export tier | cwe record contents |
|---|---|---|
No pai_internal key (any account tier, Free through Enterprise) | — | No access. Not 403-on-a-dataset: the whole surface is unreachable. |
pai_internal key, no global_intel stamp | base | The 13 public MITRE fields only. No threat_class/impact, no taxonomy, no statistics key at all. |
pai_internal key whose global_intel_tier is pro or enterprise | premium | All 13 base fields plus the 10 taxonomy fields and the full statistics block. |
This is not the Free/Registered/Pro/Enterprise ladder, and premium is not
"Enterprise-only". On the live request path the only input that lifts a caller to
premium is a global_intel stamp — a separate product entitlement.
resolve_export_tier also reads a per-key export_tier stamp and the org's
mpi entitlement, but neither is reachable today:
PAIKeyInfo has no export_tier field and nothing sets one, and
bulk_export_pai._caller_and_tier calls the resolver without an
entitlement_lookup. Both branches are unit-tested and dormant. An Enterprise account
without the global_intel entitlement resolves to base, exactly like a
Pro one. An unrecognised stamp value resolves to base — fail closed.
Because cwe is tier_required: base, it is accessible: true
in the manifest for both tiers, and a base caller never receives
403 on it — it receives the same 969 records with the premium half of each
record absent. That differs from a premium-only dataset, which stays listed with
accessible: false and returns 403 on a snapshot request.
Per-tier response. Same record count, same ordering, same ids at both tiers; only the field set differs. One NDJSON line for CWE-79.
base (artifact builds to 188,816 bytes over 969 records):
{"id":"79","name":"Improper Neutralization of Input During Web Page Generation ('Cross-site Scripting')","description":"The product does not neutralize or incorrectly neutralizes user-controllable input before it is placed in output...","extended_description":"There are many variants of cross-site scripting...","abstraction":"Base","status":"Stable","parents":[{"cwe_id":"74","view_id":"1000","is_primary":true}],"capec_ids":["209","588","591","592","63","85"],"likelihood_of_exploit":"High","owasp":{"2021":{"category_id":"1347","category_name":"OWASP Top Ten 2021 Category A03:2021 - Injection","code":"A03"}},"owasp_super_class":{"2021":"1344"},"view_1400_category":{"category_id":"1409","category_name":"Comprehensive Categorization: Injection"},"cwe_top25_views":["1200","1337","1350","1387","1425","1430","1435"]}
premium (artifact builds to 243,964 bytes over the same 969 records) — every
key above, plus:
{"threat_class":"Sensitive Information Disclosure","impact":"Cross-Site Scripting (XSS)","class_via":"master","impact_via":"leaf","class_confirmed":true,"impact_confirmed":true,"risk":"MEDIUM","software_category":null,"aggregated_category":"Cross-site Scripting (XSS)","root_cause":"cross-site scripting","statistics":{"cwe_top25_rank":1,"nvd_count":50707,"kev_count":1,"nvd_share":20.8788,"kev_share":1.5873,"epss_max":0.99293,"epss_min":0.00091,"epss_median":0.00507,"epss_mean":0.0131203,"cvss_max":10.0,"cvss_min":0.0,"cvss_mean":5.6447909,"github_poc_links":984,"by_year":null,"cwe_top25_score":null,"metasploit_modules":null,"nuclei_templates":null,"exploitdb_entries":null,"verified_exploits":null,"distinct_products":null,"distinct_vendors":null,"first_seen":null,"last_seen":null,"mean_days_to_kev":null}}
The 969/188,816/243,964 figures are a real build against the live artifacts on 2026-09-07, not an estimate.
Field tiering. The split mirrors
backend/app/services/cwe_intelligence_access.py's _BASE_FIELDS /
_PREMIUM_TAXONOMY_FIELDS, so this export and the live /cwe-intelligence
API cannot disagree about what a tier may read.
| Tier | Fields |
|---|---|
base | id, name, description, extended_description, abstraction, status, parents, capec_ids, likelihood_of_exploit, owasp, owasp_super_class, view_1400_category, cwe_top25_views — public MITRE knowledge, readable off cwe.mitre.org |
premium | everything above, plus threat_class, impact, class_via, impact_via, class_confirmed, impact_confirmed, risk, software_category, aggregated_category, root_cause, and the whole statistics block — Phoenix's own analyst work product |
The statistics PARENT is what carries the premium classification, not
its 24 individual fields. project_record filters top-level keys only and never
recurses, so registering the parent is the correct mechanism: a base caller never sees
the key at all, and a statistic added upstream later cannot leak by being forgotten in the
registry.
Coverage — these nulls are real absence, not a failed fetch. Measured against the live artifacts on 2026-09-07:
threat_class; 479 carry
an impact. The taxonomy is Phoenix analyst work product and does not cover the whole
catalogue.statistics block. The remaining
211 carry statistics: null — present and null, never omitted,
so a consumer can distinguish "no CVE activity was measured for this weakness" from "this build
dropped the field".statistics fields are always null platform-wide
because no pipeline source exists for them yet: cwe_top25_score,
metasploit_modules, nuclei_templates, exploitdb_entries,
verified_exploits, distinct_products, distinct_vendors,
first_seen, last_seen, mean_days_to_kev. A documented
follow-up, not a silent gap.Row count differs from the API on purpose.
GET /api/v1/analysis/cwe-intelligence?scope=catalog returns 989 rows;
this dataset returns 969. The 20 extra rows on the API side are CWEs MITRE has
reclassified from Weakness to Category and which therefore carry name: null and no
description. A bulk mirror keyed on identity should not carry rows whose entire body is null.
Build state. pipeline/build_bulk_exports.py is not invoked by any
scheduler — not docker/data-updater/entrypoint.sh, the Makefile,
.github/, scripts/, or pipeline/data_update_lanes.py
(verified 2026-09-07). Every bulk-export artifact is built by hand today, so cwe
appears in the manifest only after someone runs
python pipeline/build_bulk_exports.py --dataset cwe.
cve_taxonomy dataset (added 2026-09-23)tier_required: base, 314,713 records (measured 2026-09-23 against the full NVD
cache), keyed by CVE id. Built by pipeline/bulk_export/serialisers/cve_taxonomy.py; the
chain logic is pipeline/cve_taxonomy.py, shared with the public per-CVE route
GET /api/v1/cves/{cve_id}/taxonomy. The two agree only for the same source artifacts: a snapshot
reflects the artifacts at its build time, while the route reads the live ones. When the sources may have changed,
compare the manifest's source_freshness with the route's sources[*].fetched_at, and rebuild the
export.
| Step | Source |
|---|---|
| CVE → CWE | NVD weaknesses in data/cache/nvd.json; NVD-CWE-Other / NVD-CWE-noinfo dropped |
| CWE → CAPEC | MITRE CWE Related_Attack_Patterns (web/data/cwe_catalog.json), direct links only |
| CAPEC → ATT&CK | MITRE CAPEC Taxonomy_Mappings with taxonomy ATTACK (web/data/resources/capec_db.json) |
Record universe. Every non-Rejected NVD CVE with at least one real
CWE-<n>. CVEs with only NVD sentinel weaknesses, or none, are not emitted
(60,111 of 374,824 non-rejected CVEs on 2026-09-23). Rejected CVEs are not emitted either. A CVE
absent from the snapshot can therefore mean any of three things: it has no CWE classification, it is
Rejected, or it entered NVD after this snapshot was built. Read absence as "no CWE classification" only
for a CVE confirmed non-Rejected and present in the same build's source universe — check it against
cve_core from the same export generation.
Record shape — IDs only; names live once in the cwe dataset and the
/api/v1/analysis/resources/{capec-db,techniques-db} endpoints:
{"cve_id":"CVE-2000-1205","cwe_ids":["CWE-79"],"capec_ids":["CAPEC-63","CAPEC-85","CAPEC-209","CAPEC-588","CAPEC-591","CAPEC-592"],"attack_technique_ids":[],"cwe_capec":{"CWE-79":["CAPEC-63","CAPEC-85","CAPEC-209","CAPEC-588","CAPEC-591","CAPEC-592"]},"capec_attack":{"CAPEC-63":[],"CAPEC-85":[],"CAPEC-209":[],"CAPEC-588":[],"CAPEC-591":[],"CAPEC-592":[]},"attack_derivation":"structural_cwe_capec"}
cwe_capec and capec_attack keep the chain explicit, so a consumer can tell
which CWE produced a given technique.
Structural, not observed. attack_technique_ids is derived CWE → CAPEC →
ATT&CK and labelled attack_derivation: "structural_cwe_capec" on every record: techniques
possible for the weakness class (measured 0.015 micro-F1 against expert gold), not techniques observed for
the CVE. 98,603 CVEs carry at least one technique.
Per tier. Every field is public MITRE reference data and registered base, so the
base and premium artifacts are byte-identical (6,882,812 bytes compressed on 2026-09-23).
Admission is the same as every other dataset: a pai_internal key; phx_gintel_ is rejected.
| Caller | Export tier | cve_taxonomy record contents |
|---|---|---|
No pai_internal key (any account tier, Free through Enterprise) | — | No access. The whole export surface is unreachable. |
pai_internal key, no global_intel stamp | base | All seven fields |
pai_internal key with a global_intel stamp | premium | The same seven fields — byte-identical artifact |
Build. No scheduler invokes the producer. Run
python -m pipeline.build_bulk_exports --dataset cve_taxonomy (measured 87 s, 1.37 GB peak RSS).
patch_window dataset (added 2026-09-23; row shape + delta updated 2026-09-26 — R7 sharding)tier_required: premium, one row per CVE that has a patch-window record, keyed by
CVE id. Built by pipeline/bulk_export/serialisers/patch_window.py, following the
cve_taxonomy.py template. Source is the sharded generation under
web/data/patch_windows/ (pipeline/patch_windows/shards.py, 4096 shards,
generation-versioned behind a current.json pointer — see the R7 sharding note in
2026-09-23-patch-window-remediation.md),
streamed one shard at a time (shards.iter_records) — never the old single ~2.48 GB
by_cve.json, which no longer exists.
Row order is shard order, not sorted by CVE id. shards.iter_records walks the
4096 shard files in filename order and yields each shard's records in on-disk dict order; a consumer that
needs sorted output must sort client-side.
Record shape no longer carries patch_window_generated_at or
patch_window_sources. Both are dataset-level facts (identical on every row of one build)
that previously (a) marked every one of ~64k rows "updated" on every delta even when nothing about
that CVE changed, and (b) changed the snapshot's bytes on every rebuild, defeating the ETag/304
contract. They now live only in the manifest's per-dataset source_freshness block (above), keyed
patch_window_generated_at / patch_window_sources; every API surface that serves the
unshaped record (CVE detail, batch, PAI mirror, MCP) is unaffected — this exclusion is
bulk-export-only.
{"cve_id": "CVE-2026-1234", "patch_window": {"patch_window_status": "contained", "patch_window_summary": "...", "patch_window_vendors": [...]}}
Delta-enabled (delta_enabled=True, delta_retention=30) — see
GET /internal/v1/export/patch_window/delta below. Retention is 30 builds:
since older than the retention floor gets 409 delta_cursor_expired, not a
partial/truncated delta.
ETag stability. Because the two build-scoped keys above no longer sit in the row, a rebuild
that changes no record's patch_window content produces byte-identical snapshot bytes and therefore
the same ETag / sha256_uncompressed as the previous build —
If-None-Match returns 304 across a no-op rebuild, which was not previously possible for
this dataset.
| Caller | Export tier | patch_window record contents |
|---|---|---|
No pai_internal key (any account tier, Free through Enterprise) | — | No access. The whole export surface is unreachable. |
pai_internal key, no global_intel stamp | base | No access to this dataset. Its DATASET_MIN_TIER is premium, so a base-tier caller never gets a build for it at all. |
pai_internal key with a global_intel stamp | premium | cve_id plus the full patch_window object |
Nested-key safety. Registered in tests/pipeline/test_bulk_export_nested_field_safety.py,
which enforces that no premium-classified field name anywhere in EXPORT_PROJECTIONS
collides with a same-named field another dataset serves as base-safe raw evidence — the same
class of defect the CRA/SBOM library-intelligence work found and fixed by renaming rather than tolerating.
Build. No scheduler invokes the producer yet. Run
python -m pipeline.build_bulk_exports --dataset patch_window.
cve_intel dataset (added 2026-10-01, ASPMAIN-7799)tier_required: base, one row per CVE, keyed by CVE id, sorted by CVE id. Built by
pipeline/bulk_export/serialisers/cve_intel.py. Gated by its own flag
enable_pai_cve_intel_dataset (default false,
DatasetSpec.required_flag, resolved through feature_status.resolve_flag) on top of
enable_pai_bulk_export: when it is off the dataset is absent from the manifest and both
/internal/v1/export/cve_intel/snapshot and /internal/v1/export/cve_intel/delta return
404 (never 403). The producer builds the artifact regardless of the flag.
Rows are the UNION, not the intersection, of three membership sources: PS-HP
(data/internal/high_profile_intelligence_internal.json, legacy web/data copy only
when the internal file is absent — high_profile[] and enterprise_watchlist[]), bug bounty
(bugbounty_intelligence.json cve_lookup) and GitHub PoCs
(github_exploits.json). Roughly 15k-19k CVEs. A block is null when the CVE is absent
from that block's source(s).
Absent vs corrupt sources. A source file that does not exist is skipped: its blocks are
null for every row. A source that is present but unreadable, not valid JSON, or of the wrong
shape (wrong top-level type, or a section such as cve_lookup or high_profile with the
wrong type) fails the cve_intel build: nothing is published, the previous latest.json
and delta files stay as they were, and the producer run exits 1. A present but corrupt internal PS-HP file
does not fall back to the legacy copy. A single malformed value inside a valid file is skipped and counted
in one WARNING.
News enriches, it never adds rows. The newest
news_tracking/daily/news_<date>.json listed in that directory's manifest.json
(all_mentions) only fills popularity.news for CVEs already in the set above.
CVEs that appear only in the rolling news cohort are not in this dataset — use
POST /internal/v1/cves/batch with include=popularity for them. The cohort swings by
hundreds of CVEs day to day; as membership it would delete far more rows than the delta delete valve allows
(0.5%) and block the dataset until a manual override. When a member CVE leaves the news window its row is an
ordinary delta upsert (popularity.news reverts to the unlisted values, or
popularity becomes null if the CVE has no bug-bounty entry either), never a
delete.
{"cve_id": "CVE-2024-0001",
"ps_hp": {"status": "high_profile", "ps_hp_score": 88.5, "ps_hp_tier": 1, "ps_hp_tier_name": "Confirmed",
"components": {...}, "hp_reasons": [...], "hp_summary": "...", "hp_rationale": "...",
"enterprise_category": "network", "is_enterprise_watchlist": false},
"popularity": {"bug_bounty": {"listed": true, "total_reports": 12, "best_rank": 3, "popularity_score": 40.2,
"appearances": 2, "is_popular": true},
"news": {"hot": true, "mention_count": 3, "unique_sources": 2, "score": 10,
"first_seen": "...", "window_date": "2026-09-30", "top_articles": [...]}},
"github_pocs": {"poc_count": 4, "poc_level": "high", "has_github_poc_high": true, "popularity_score": 0.42,
"total_stars": 120, "total_forks": 30, "is_aggregate": true},
"exploitation_tier": "WEAPONIZED"}
Same field names and shapes as the live batch path. ps_hp,
popularity and github_pocs are field-for-field the include=ps_hp /
popularity / github_pocs blocks of POST /internal/v1/cves/batch, and
exploitation_tier is that endpoint's intel.exploitation.exploitation_tier (the platform
ladder canonical_tier_from_flags: WEAPONIZED, ACTIVELY_EXPLOITED,
EXPLOIT_AVAILABLE, THEORETICAL).
tests/pipeline/bulk_export/test_cve_intel_parity.py runs both on one fixture and asserts equality.
Differences, all deliberate:
ps_hp is null for a CVE in neither PS-HP list, where the batch path returns
{"status": "not_high_profile", ...nulls}. popularity is null when the CVE is
in neither the bug-bounty nor the news source; when it is in one, the other half carries the batch path's
"unlisted" values. github_pocs is null unless the CVE is in the PS-HP
high_profile list or has a GitHub PoC entry.github_pocs also reads github_exploits.json's cve_lookup map, through the
same aggregate formula. The live batch path reads only top-level CVE keys and the 100-row
top_by_repos, so for most of the ~7.9k cve_lookup CVEs it reports
poc_count: 0 today; the bulk row carries the real count. A top-level CVE key holding an
empty list is treated as absent (no row of its own, no poc_count: 0 block, and it does not
hide a top_by_repos/cve_lookup aggregate for that CVE); the batch path reports
poc_count: 0 for it.exploitation_tier is null only when every exploitation source
(exploit_database.json, github_exploits.json,
sightings_exploit_verification.json) is missing — the batch path's
source_unavailable case. The exploitation_evidence dataset still does not emit this
field.Delta-enabled (delta_enabled=True, delta_retention=30, same as
patch_window) — see GET /internal/v1/export/cve_intel/delta below.
github_pocs.total_stars, github_pocs.total_forks and
github_pocs.popularity_score are excluded from the per-record delta hash (DatasetSpec.hash_exclude_fields): they stay in every snapshot row and
in every delta upsert body, but a change confined to them produces no delta op. A
consumer that applies only deltas therefore holds star/fork counts and the star-derived
popularity_score as of the last real change to that CVE (a
PS-HP tier, watchlist mark, popularity, PoC count or exploitation-tier change); re-baseline from the snapshot if
exact counts matter. popularity_score is excluded because, for a CVE whose PoC summary is not
PS-HP-matched, it is a continuous function of stars/forks. poc_level is NOT excluded: a threshold
crossing (low/medium/high) is a real change and produces a delta op.
For the ~382 PS-HP CVEs, ps_hp.components.popularity and ps_hp.ps_hp_score are also
star-influenced and still count in the delta hash, so expect a few hundred upsert ops per day
from star movement alone.
Per tier. Every field is base (EXPORT_PROJECTIONS["cve_intel"]),
matching the un-gated exposure of the same blocks on POST /internal/v1/cves/batch; a
premium caller gets the same row. source_freshness reports each source file's mtime
(<source>_generated_at) plus news_window_date.
Build. python -m pipeline.build_bulk_exports --dataset cve_intel.
product_health dataset (added 2026-10-01, ASPMAIN-7800)tier_required: premium, snapshot-only ("delta": {"supported": false};
there is no /internal/v1/export/product_health/delta route). One row per product in
data/internal/product_health_index.json (~121,558 products), keyed by
{vendor}:{product} (the index key: NVD CPE vendor and product, lowercased with
spaces as _, exactly as POST /internal/v1/products/health normalises its input),
sorted by that key. Built by pipeline/bulk_export/serialisers/product_health.py. Pull it with
GET /internal/v1/export/product_health/snapshot.
Gated by enable_pai_product_health — the same flag as
POST /internal/v1/products/health (DatasetSpec.required_flag, resolved through
feature_status.resolve_flag), on top of enable_pai_bulk_export. When it is off the
dataset is absent from the manifest and its snapshot route returns 404 (never 403),
even for a premium caller. The producer builds the artifact regardless of the flag.
Requires the data/internal mount. The index lives outside the nginx root in
data/internal/ (written by pipeline/product_health_index.py). In Compose,
/app/data/internal is the named volume internal-data, mounted read-write in the
data-updater and read-only in the backend; the git-tracked ./data/internal is only the
read-only seed, mounted in the data-updater at /app/data/internal-seed and copied into the
volume on start for files it lacks (ASPMAIN-7798). With no index file at all
the dataset is skipped (product_health:premium in the build's skipped_absent list,
one WARNING naming the path) and the run does not fail. A present but unreadable or malformed
index fails the build for this dataset only (product_health:premium in failed, an
ERROR log line). Either way nothing is published, any previously published artifact stays
current, and an empty dataset is never published.
Same record as the live route. Each row is exactly the per-product item
POST /internal/v1/products/health returns for that pair, minus the request-matching
status (always "ok" here) and match (its key is the row key):
{"vendor": "microsoft", "product": "windows_10",
"health": {"phs": 41.2, "prs": 58.8, "grade": "D", "domain_scores": {...}, "buckets": {...},
"inputs": {"cve_count": 100, "...": "the stored counts"}},
"fix": {"fix_available_rate": 0.62, "unfixed_count": 38, "unfixed_kev_count": 1,
"evidence": "nvd_reference_tags"},
"watchlist": {"is_enterprise_watchlist": true, "cve_ids": ["CVE-2024-0003", "CVE-2024-0007"]}}
fix is null when the product has no CVEs (cve_count 0), as on the live
route. Scores are computed at build time, not stored: health comes from
calculate_product_score(counts, None) (pipeline/scoring/phs_scoring.py) and the
blocks and watchlist join from pipeline/product_health_lookup.py — the same functions the
live route calls (the backend re-exports them).
tests/pipeline/bulk_export/test_product_health_serialiser.py asserts every row equals the live
lookup() output on the same files.
Watchlist source and freshness. watchlist.cve_ids comes from the PS-HP internal
file's enterprise_watchlist (data/internal/high_profile_intelligence_internal.json;
the legacy web/data copy only when the internal one is missing, the live route's
order). With neither file, every watchlist is empty, as on the live route, and the build logs a
WARNING. A PS-HP file that is present but unreadable or not a JSON object fails the
product_health build (nothing published) instead of publishing empty watchlists; the live
route falls through to the next file per request instead. source_freshness reports the index's own generated_at /
nvd_generated_at (product_health_index_generated_at,
product_health_index_nvd_generated_at) and the PS-HP file mtimes in the cve_intel
style (ps_hp_internal_generated_at, ps_hp_legacy_generated_at). Rows reflect the
index and PS-HP file as of the build, so they can lag the live route by up to one build cycle.
Per tier. Every field is premium (EXPORT_PROJECTIONS["product_health"]),
matching the route's PAI-only posture. No base artifact is built; a base caller sees
the dataset listed with accessible: false and gets 403 on the snapshot.
Snapshot-only on purpose. About 121k small records (low tens of MB) are cheap to re-pull on every build. Delta will be reconsidered if consumer sync time says otherwise.
Build. python -m pipeline.build_bulk_exports --dataset product_health.
product_vendor dataset (added 2026-10-01, ASPMAIN-7801)This dataset is NOT complete. Ambiguous products are NOT in it. It holds only products for which exactly one vendor in the index has that product name. A product with two or more candidate vendors (for example
openssh, owned by bothopenbsdandopenssh) is omitted entirely — never exported with a guessed or “first” vendor. A product name missing from this dataset MUST be resolved viaPOST /internal/v1/products/resolve-vendor; do not treat absence as “unknown product” or “no vendor”. That route returnsstatus: "ambiguous"with rankedcandidates, narrows byvendor_hint/cve_ids, or reportsnot_found.
tier_required: base, snapshot-only ("delta": {"supported": false};
there is no /internal/v1/export/product_vendor/delta route). One row per unambiguous product in
data/internal/product_health_index.json, keyed by the product name (lowercased,
spaces as _, as resolve-vendor normalises its input), sorted by that key. Built by
pipeline/bulk_export/serialisers/product_vendor.py. Pull it with
GET /internal/v1/export/product_vendor/snapshot.
Row shape — the same values resolve-vendor returns for an ok /
matched_by: "unique_product" item:
{"product": "windows", "vendor": "microsoft", "cpe_prefix": "cpe:2.3:o:microsoft:windows"}
cpe_prefix is cpe:2.3:<cpe_part>:<vendor>:<product> with the index row's
real cpe_part (a/o/h, default a).
“Unambiguous” is exactly live rule 4 (“exactly one vendor has this
product”), computed with the same vendors_by_product reverse map and
cpe_prefix builder as the route (pipeline/product_health_lookup.py; the backend
re-exports them). The other rules cannot be precomputed, so they are not in this dataset: rule 1
(invalid product names such as none/null are skipped), rules 2-3 (depend on
the caller's per-request vendor_hint), rule 5 (needs cve_ids and a live NVD CPE lookup),
rule 6 (ambiguous — no single vendor to export) and rule 7 (not_found).
tests/pipeline/bulk_export/test_product_vendor_serialiser.py asserts every row equals the live
resolve_vendor result and that ambiguous product names are absent.
Intended use. A nightly sync can pre-resolve every unambiguous product to its vendor; only the
products absent from this dataset need the live resolve-vendor batch route.
Gated by enable_pai_product_resolve_vendor — the same flag as
POST /internal/v1/products/resolve-vendor (DatasetSpec.required_flag, resolved through
feature_status.resolve_flag), on top of enable_pai_bulk_export. Off: absent from the
manifest and the snapshot route returns 404 (never 403). Requires the
data/internal mount; with no index file the dataset is skipped
(product_vendor:base and product_vendor:premium in the build's
skipped_absent list, one WARNING) without failing the run; with a present but
malformed index the build fails for this dataset only (both tiers in failed, an
ERROR log line). Either way it publishes nothing. source_freshness reports
product_health_index_generated_at and product_health_index_nvd_generated_at.
Per tier. Every field is base (EXPORT_PROJECTIONS["product_vendor"]):
public NVD vendor/product naming only, no scores, counts or tenant data. premium callers receive the
same rows.
Build. python -m pipeline.build_bulk_exports --dataset product_vendor.
phoenix_scores, malware_packages, library_license (added 2026-10-01, ASPMAIN-7802)Three tier_required: premium datasets carved out of the same
web/data/library_intelligence.json package records that library_intel exports. Each holds
fields library_intel deliberately never exports, at any tier. All three are keyed by the package
purl (the same id as library_intel, so rows join one-to-one), have every
field premium (EXPORT_PROJECTIONS), and are gated by enable_pai_bulk_export
plus the premium tier only: no per-dataset required_flag. A base caller sees
them listed with accessible: false and gets 403 on the snapshot (and on the
malware_packages delta). No base artifact is built.
Fields moved here, not in library_intel. library_intel's
base allow-list is unchanged (its output and hashes are byte-identical before and after these
datasets were added):
Source field(s) in library_intelligence.json | Bulk dataset | Notes |
|---|---|---|
pss | phoenix_scores | Phoenix's PS-OSS package score, passed through verbatim |
compromised, compromised_status, compromised_versions, compromised_safe_versions, compromised_campaigns, compromised_severity, compromised_last_update, compromise_intel | malware_packages | Compromised packages only |
license, license_category, license_risk_score, license_policy | library_license | R7: licence data stays premium |
license_osi_approved, license_source, license_policy_decision, license_risk_level | not exported | Not in the planned shape; still dropped everywhere |
phoenix_scores — snapshot-only (no
/internal/v1/export/phoenix_scores/delta route). One row per package that has a pss
block, sorted by purl:
{"purl": "pkg:npm/lodash",
"pss": {"total": 90, "components": {"blast_radius": 0.7, "compromise": 1.0, "...": "..."},
"floor_applied": "compromised_active", "floor_value": 90,
"exploitation_evidence": {"tier": "SUPPLY_CHAIN_COMPROMISED", "flags": {...},
"github_stats": {...}, "counts": {...}}}}
pss is the producer's stored block (pipeline/rebuild_unified_library_intel.py), never
recomputed: pss.total is the same stored value that
POST /internal/v1/intel/library_versions/expanded surfaces as exploitation.pss_score.
Packages without a pss block are omitted. Built by
pipeline/bulk_export/serialisers/phoenix_scores.py.
PS-OSS package scores only; not the CVE-scale PS-HP/PS-PHS/PS-TTS outputs that the 2026-08-19 bulk-export design reserved the phoenix_scores name for.
malware_packages — delta-enabled
(delta_retention=30, like patch_window and cve_intel). A new compromise mark is
a rare but urgent event, so a consumer can catch it within one build with
GET /internal/v1/export/malware_packages/delta (see the delta endpoint below) instead of
re-pulling the snapshot. Row shape:
{"purl": "pkg:npm/@emilgroup/public-api-sdk", "compromised": true,
"compromised_status": "compromised", "compromised_versions": ["1.33.1", "1.33.2"],
"compromised_safe_versions": [],
"compromised_campaigns": [{"name": "EMILGROUP / teale.io npm malware campaign",
"date": "2026-03-21", "severity": "CRITICAL", "threat_actor": "TeamPCP",
"tags": ["canister-worm"]}],
"compromised_severity": "CRITICAL", "compromised_last_update": "2026-03-23",
"compromise_intel": {"compromised": true, "days_since_last_incident": 16,
"incident_count_365d": 1, "repeat_offender": false, "ecosystem_wide": true,
"ci_injection": false, "decay_phase": "30d", "compromise_score": 1.0,
"exploitation_tier": "SUPPLY_CHAIN_COMPROMISED"}}
Only packages with compromise data are in it: compromised true, a
compromised_status other than none (for example potential), non-empty
compromised_versions or compromised_campaigns, or a positive compromise_intel
signal (compromised, incident_count_365d > 0, repeat_offender,
compromise_score > 0). Absence means "not flagged as compromised", not "unknown package". When a
package is newly flagged it arrives as a new upsert in the delta. When the source stops flagging it,
it arrives as a delete. The day counter compromise_intel.days_since_last_incident stays in
the row but is not in the per-record hash (hash_exclude_fields), so its daily tick
does not mark every row updated. Built by
pipeline/bulk_export/serialisers/malware_packages.py.
Build safety guard (small dataset). Small retractions publish: up to 10 un-flagged packages
per build (DatasetSpec.delete_floor=10) and a net drop of up to half the published rows
(max_drop_fraction=0.5) go out without an override, so new compromise marks are not held back. A
drop of more than half, or a wipe (zero rows or a missing source), is refused until an operator override
(PHOENIX_EXPORT_ALLOW_MASS_DELETE / PHOENIX_EXPORT_ALLOW_SHRINK); the previous build
stays current. Every other dataset keeps the default guards (max(1, floor(0.5%)) deletions, 10%
net drop).
This dataset carries the compromise fields library_intelligence.json holds per package. It is
not a dump of the MPI scan corpus (compromised_package_intel / verdict reasoning /
IOCs).
library_license — snapshot-only (no
/internal/v1/export/library_license/delta route). Row shape, always all four keys
(null where the source has none):
{"purl": "pkg:npm/lodash", "license": "MIT", "license_category": "permissive",
"license_risk_score": 0.1, "license_policy": "allow"}
"UNKNOWN" / "deny" rows are real licence data and are kept. A package with none of the
four fields is omitted. Built by pipeline/bulk_export/serialisers/library_license.py.
Shared source and freshness. All three read library_intel's cached parse of the
file, so one build parses it once for all four datasets. source_freshness and
source_fingerprint are library_intel's
({"generated_at": <_metadata.generated_at>}). A value of an unexpected type (for example a
non-object pss or compromise_intel) is skipped or nulled with one counted
WARNING per build, never a crash. A missing source yields zero rows, which the builder refuses to
publish (any prior artifact stays current).
Build. python -m pipeline.build_bulk_exports --dataset phoenix_scores (likewise
malware_packages, library_license).
GET /internal/v1/export/{dataset}/snapshotStreams the full NDJSON snapshot for one dataset at the caller's resolved tier, gzip-compressed
on disk and served with Content-Encoding: gzip (the body is not re-compressed
by the app-wide GZipMiddleware — the pre-set Content-Encoding header
short-circuits it).
Response headers:
| Header | Meaning |
|---|---|
ETag | "<sha256>" of the compressed artifact; use with If-None-Match. See "Verifying integrity" below before hashing anything against it |
X-Phoenix-Export-Sha256-Uncompressed | sha256 of the decompressed NDJSON — the digest a normal HTTP client can reproduce. Empty string on artifacts published before this field existed; treat "" as "not published", never as a mismatch |
X-Phoenix-Export-Cursor | Opaque, resumable position (base64url; see pipeline.bulk_export.cursors). Persist this — it is what a consumer presents as from on the first delta once Phase 3 ships. Decodes (server-side only) to (dataset, tier, seq); never parse it client-side, it is not a security boundary but is not a stable format either. |
X-Phoenix-Export-Snapshot-Id | Human-readable build label, {dataset}-{tier}-{seq:06d} (e.g. kev-base-000412). Convenient for logs/support tickets; not a substitute for X-Phoenix-Export-Cursor. |
X-Phoenix-Export-Watermark | ISO-8601 timestamp this snapshot was built as-of. Informational only — never use this as a resume position or range key; it is exactly the wall-clock range-key pattern R2 rules out (silently drops records under concurrent ingest or clock skew, indistinguishably from "nothing changed"). |
X-Phoenix-Export-Dataset | Echo of the requested dataset name |
X-Phoenix-Export-Tier | Tier the artifact was resolved at (base / premium) |
X-Phoenix-Export-Records | Row count in the artifact |
X-Phoenix-Export-Schema-Version | 1.0 |
Cache-Control | private, max-age=0, must-revalidate — the artifact is per-tier and per-caller, so it must not be stored by a shared cache; revalidate with If-None-Match against the ETag rather than serving a stale copy |
Last-Modified | The artifact file's on-disk mtime, RFC 1123-formatted (email.utils.formatdate) |
Content-Disposition | attachment; filename="<snapshot file>" |
Conditional requests. Send If-None-Match: "<etag>"; a match returns
304 with an empty body and no re-transfer. Only If-None-Match is honoured
— Last-Modified is informational and If-Modified-Since is
not implemented.
Verifying integrity — read this before comparing digests. Two digests are published because no single one is reproducible by every client.
| How you read the body | Digest to compare against |
|---|---|
Normal HTTP client (requests, httpx, Go, curl --compressed) — the library transparently decompresses | X-Phoenix-Export-Sha256-Uncompressed |
Raw undecoded bytes (resp.raw.read(decode_content=False), curl without --compressed, writing straight to a .gz file) | ETag / the manifest's sha256 |
The response carries Content-Encoding: gzip, so a standard client decompresses before
your code sees the bytes, and hashing what you received can never reproduce the ETag
(which covers the compressed artifact). Comparing the wrong pair looks exactly like corruption. Both
digests are watermark-independent: an unchanged dataset rebuilt tomorrow yields identical values,
which is what makes 304 revalidation worthwhile.
Range requests. Lets a consumer resume a dropped multi-gigabyte transfer instead of
restarting it. Support comes from the installed Starlette's FileResponse, and was
verified live against the deployed runtime (Starlette 1.6.0): a
Range: bytes=0-99 request returned 206 with exactly 100 bytes and a
Content-Range header. Range handling was added to FileResponse well before
1.0, so the exact earliest supporting release is deliberately NOT stated here — only the version
actually tested is. backend/requirements.txt pins starlette>=0.27.0, which
does permit resolving a version older than the one verified, so an unchecked deployment must be treated
as "may not support it": a Range request answered 200 with no
Content-Range/Accept-Ranges has no range handling. If
you must support both, treat 206 as the fast path and fall back to a full re-fetch when the
response is 200.
Status codes:
| Condition | Code | Rationale |
|---|---|---|
| Flag off | 404 | House rule — never leak feature existence, same as everywhere else in PAI |
No/invalid key, wrong scope (incl. phx_gintel_) | 401 / 403 | Standard PAI auth |
| Unknown dataset (not in the manifest and not a registered pipeline dataset) | 404 | |
| Dataset registered/listed but requires a higher tier than the caller holds | 403 | Deliberate divergence from the 404 default — the manifest (or the pipeline registry) already told this caller the dataset exists, so a 404 here would be actively confusing |
Export generation not built yet (no manifest.json), or this dataset not built for the caller's tier | 503 + Retry-After: 3600 | Honest — not an empty 200, which would read as "dataset has no rows" |
| Second concurrent/rapid-reissue request for the same key | 429 | See concurrency guard below |
| Concurrency lease backend unreachable (cache outage) | 503 | Fails closed on this route only — see concurrency guard below for why |
| Match | 200 (or 304 on If-None-Match hit) |
Metering. /internal is exempt from the global usage-tracking middleware
(main.py), so this router meters itself: every snapshot request (200 or 304 — the check
runs before the conditional-request branch) increments the caller's exports counter one time
via UserTierService().record_usage(subject, {"exports": 1}) — the same rail
user_tier_service.FEATURE_USAGE_FIELDS already maps the bulk_export feature to. The
subject is the key's admin_id (falling back to key_id), matching every other PAI
usage-metering call site (intel_query_pai._usage_subject). GET /manifest records
nothing — discovery stays free. A metering failure (e.g. storage unavailable) is logged and
swallowed; it never fails the export request itself.
Per-key concurrency guard. PAI's standard rate limit (1000 calls/minute,
config.py) is the wrong shape for full-corpus streams — one consumer looping a
multi-gigabyte snapshot download would saturate the backend. This router additionally caps each key
to EXPORT_CONCURRENT_STREAMS_PER_KEY (currently 1) snapshot streams in flight at a
time, tracked as a cache-backed counter (CacheService.increment /
CacheService.increment_with_expire / CacheService.decrement) that is incremented
when a stream starts and released once the response has actually finished sending. A second request
for the same key while one is still in flight gets 429 Export already in progress for this
key.
The guard fails closed on a cache outage: if a concurrency lease cannot be
established authoritatively, the snapshot route returns 503 Export concurrency control
unavailable; retry shortly. This is a deliberate divergence from the fail-open convention
every interactive PAI route follows, and it is scoped to this one streaming route —
GET /internal/v1/export/manifest still serves normally during an outage, so a consumer
can always discover state and back off.
The reasoning: PAI's ordinary per-key rate limiter depends on the same cache, and the
exports counter above is accounting-only (this router calls no
UserTierService.enforce_feature_limit against it), so failing open did not degrade one
control — it removed every control at once, on the only route that streams
multi-hundred-megabyte to multi-gigabyte bodies. One valid or stolen key could then open unlimited
concurrent streams and exhaust disk and egress capacity precisely while the platform was already
degraded. Refusing bulk egress that cannot be accounted for is the correct trade here. Note this
remains a per-key concurrency cap, not a volume budget: no per-day export quota is
enforced in this phase.
GET /internal/v1/export/{dataset}/delta (added 2026-09-26, Phase 3)Serves the changed-since feed for a delta-enabled dataset, as multi-member gzip NDJSON.
backend/app/routers/bulk_export_delta_pai.py (a separate module from
bulk_export_pai.py, which is at the file-size limit; every shared helper it calls is called
through the bulk_export_pai module object, never imported by name, so a test
monkeypatching that module reaches this route too).
Gates, in order:
enable_pai_bulk_export off → 404 (the whole surface).required_flag (patch_window →
enable_patch_windows, cve_intel → enable_pai_cve_intel_dataset,
product_health → enable_pai_product_health, product_vendor →
enable_pai_product_resolve_vendor) off → 404. Checked before the dataset lookup below
— the same code path an unknown dataset name would also hit, since
_dataset_flag_enabled is a no-op (returns True) for a dataset with no
required_flag registered, including one the registry doesn't know at all.DatasetSpec.delta_enabled is False →
404. This includes every dataset except patch_window, cve_intel and malware_packages today — there is no
partial/degraded delta response for a non-delta dataset, only 404, identical to an unknown
dataset name.DATASET_MIN_TIER → 403.PAI_INTERNAL scope only — same admission as every other bulk-export route;
phx_gintel_ is rejected here too.Query parameters:
| Param | Required | Meaning |
|---|---|---|
since | yes | An export cursor (from a snapshot's X-Phoenix-Export-Cursor, the manifest's snapshot.cursor/delta.latest_cursor/delta.retention_cursor, or a previous delta response's X-Phoenix-Export-Cursor) naming the position already applied. Not a timestamp — cursors are opaque monotonic per-(dataset, tier) build sequences (spec R2), never wall-clock. |
until | no | A cursor naming the position to stop at (inclusive). Defaults to the dataset's latest build. Must satisfy since <= until <= latest. |
The window is half-open: (since, until]. The record at since
itself is never re-sent — only builds strictly after it, up to and including
until.
Response 200: Content-Type: application/x-ndjson. When the window
covers at least one build, Content-Encoding: gzip and the body is a multi-member gzip
stream — one gzip member per covered build's delta file, concatenated as-is (no
re-compression, no member merging).
gzip/zlib's own multi-member semantics, and most language HTTP
stacks were never built to expect it on a single response. Read the raw, undecoded bytes and loop
zlib.decompressobj(16 + zlib.MAX_WBITS) over each decompressor's unused_data until
unused_data is empty, decoding one member per iteration. A client that decodes only the first
member will silently miss every later build in the window with no error.
When since already equals the latest build (no builds in the window), the response is
200 with an empty body and no Content-Encoding header — there
is no gzip member to send, so claiming Content-Encoding: gzip on zero bytes would itself be a
malformed gzip stream.
Each decoded NDJSON line is one op:
{"op": "upsert", "id": "CVE-2026-1234", "reason": "new", "record": {"cve_id": "CVE-2026-1234", "patch_window": {...}}}
{"op": "upsert", "id": "CVE-2026-1234", "reason": "updated", "record": {"cve_id": "CVE-2026-1234", "patch_window": {...}}}
{"op": "delete", "id": "CVE-2026-1234", "reason": "withdrawn"}
reason is exactly one of new, updated, withdrawn.
No other value is ever emitted.record (upsert only) is byte-identical to that id's row in the snapshot at
the same projection — same tier, same field set, same patch_window shape documented
above.new then updated, or updated then
withdrawn) — only the final op for a given id reflects that id's state at
until.Response headers (delta and snapshot share these names; a delta response has no
X-Phoenix-Export-Records/ETag, since a delta is not a cacheable single
artifact):
| Header | Meaning |
|---|---|
X-Phoenix-Export-Cursor | The position after applying this response — store this as the next request's since |
X-Phoenix-Export-Watermark | ISO-8601 timestamp of the until build. Informational only, same caveat as the snapshot header |
X-Phoenix-Export-Dataset | Echo of the requested dataset |
X-Phoenix-Export-Tier | Resolved export tier |
X-Phoenix-Export-Schema-Version | 1.0 |
Errors:
| Condition | Code | Body |
|---|---|---|
since/until is not a valid cursor | 400 | {"detail": "...", "code": "invalid_cursor"} |
since/until was issued for a different dataset or tier | 400 | {"detail": "...", "code": "cursor_mismatch"} |
since > until, until > latest, or since > latest | 400 | {"detail": "...", "code": "invalid_window"} |
Window spans more builds than delta.max_window_seqs (the dataset's delta_retention) | 400 | {"detail": "...", "code": "window_too_wide"} |
Flag off, dataset not delta-enabled, or dataset's own required_flag off | 404 | — |
| Caller's tier does not satisfy the dataset's minimum | 403 | — |
since is older than the retention floor, OR a delta file inside an otherwise-servable window is missing from disk (already pruned) | 409 | {"detail": "...", "code": "delta_cursor_expired", "retention_cursor": "<cursor>"} — re-baseline: pull the current snapshot and resume from its X-Phoenix-Export-Cursor |
| Second concurrent/rapid-reissue request for the same key | 429 | Same per-key concurrency lease as the snapshot route |
| Concurrency lease backend unreachable, or this export generation is not built yet | 503 (Retry-After: 3600 for the unbuilt case) | |
| Match | 200 | See above |
Metering. A delta request costs 0 exports — the exports
usage counter is a snapshot-only rail. It still takes the same per-key concurrency lease as a snapshot
stream (one concurrent stream per key, shared between the two routes), so a snapshot and a delta pull for
the same key cannot run at once.
Datasets are built independently by separate producer runs and carry independent watermarks.
A kev snapshot may reference a CVE not yet present in a cve_core snapshot; a
malware record may reference a package version absent from a base dataset that has not rebuilt yet.
Consumers must tolerate dangling cross-dataset references. Do not build a foreign-key constraint across two mirrored datasets — it will hold in testing and fail in production.
This is a deliberate design trade-off, not an oversight: enforcing cross-dataset consistency would require a global build barrier across every producer, serialising the pipeline. The contract is eventual convergence — any dangling reference resolves within one build cycle of the referenced dataset.
Related: docs/plans/2026-08-19-pai-bulk-export-sync-design.md (design,
§7-§9), docs/plans/2026-08-19-pai-bulk-export-phase1-implementation.md
(implementation plan).
Flag: enable_org_team_admin_scoping (default false). All endpoints in this section return 404 when the flag is off.
Auth: Global admin or org manager (is_org_manager Cognito group) for the target org — enforced by resolve_admin_scope + ensure_target_org_in_scope.
Related: docs/Individual_Feature/2026-06-29-org-team-tenancy-plan-3-tenancy-router.md
| Endpoint | Auth | Purpose |
|---|---|---|
GET /internal/v1/admin/organizations | global-admin or org-manager | List organizations. Non-global admins see only their own org. Response also carries tier_default_seats (canonical scf_seats_{tier} map: registered 3 / pro 25 / enterprise 100; still used for vuln_intel and mpi seat defaults only — global_intel has no tier default and no floor, see the bundle licence count note under POST .../license) global_intel_enabled (enable_global_intel_license, the gate on granting a bundle) and global_intel_lift_enabled (the SEPARATE cumulative-lift gate, requiring both enable_global_intel_license AND enable_mpi_tier_separation). These two can disagree: with the licence flag on and MPI tier separation off, an org can be granted a bundle and rendered as licensed while every member still gets 403 on mint — clients must surface that mismatch. All three advisory fields are nullable (added 2026-08-18): a transient settings/flag lookup failure degrades the field to null rather than 500ing the listing; treat null as "unknown", never as false or an empty seat map, and do not auto-apply seat defaults you could not read. The write paths still fail loudly. Each org row gains global_intel_tier / global_intel_seats (null when the org holds no bundle row) — added 2026-08-17. |
GET /internal/v1/admin/organizations (cont.) | global-admin only for these fields | Added 2026-08-19. The response also carries unlinked_organizations — organization NAMES present in users.organization, pending_registrations.organization, scf_org_license.org_key or api_keys.org_id that have no scf_tenants row, so they cannot appear in organizations at all. Such an org can hold an SCF seat licence and issue API keys while being invisible to org administration. Each row is {org_key, user_count, pending_count, license_count, api_key_count}. Global admins only (the scan is deployment-wide and has no tenant to scope to); a non-global admin gets null, not a filtered list. Not gated on enable_org_adopt_and_merge — reporting that the org list is incomplete is not a feature. Linked orgs are subtracted server-side via normalize_org_slug (NFKC → casefold → confusable skeleton), so clients must not re-derive it with a plain string compare. unlinked_truncated is true when the 500-row candidate ceiling was hit; null means unavailable, which is not the same as false. adopt_merge_enabled reports whether the two endpoints below exist for this caller. |
POST /internal/v1/admin/organizations/adopt | global-admin + enable_org_adopt_and_merge | Added 2026-08-19. Promotes a free-text organization NAME into a real org: tenant, metadata, default team, vuln_intel/mpi entitlements, links any orphan scf_org_license row, and adds matching users as members. 404 (not 403) when the flag is off. The name is stored RAW as scf_tenants.org_key, never replaced by the normalized slug: SCF enrollment matches that column exactly, with no case or whitespace folding, so storing the slug would 403 every key already issued under the original spelling. Additive — users.primary_tenant_id is set only where NULL, so a user already in another org keeps that membership. 400 unnormalizable name or a seat count below the tier floor (the message names the tier and floor); 409 ORG_KEY_TAKEN / SLUG_TAKEN, including on a concurrent race. |
POST /internal/v1/admin/organizations/{org_id}/merge | global-admin + enable_org_adopt_and_merge | Added 2026-08-19. Irreversible, no dry-run, no un-merge. Merges source_org_id INTO {org_id} — the path org survives. confirm_org_key must echo the source's RAW org_key exactly (case-sensitive); if that key cannot be resolved the endpoint returns 503 and changes nothing rather than falling back to the slug. Moves every tenant_id-bearing table (discovered from pg_catalog at runtime) plus the raw org-name columns, then deletes the source — all in one transaction, audited as org_merge in audit_log within it. scf_org_license is keyed by org_key STRING, so the source's licence row is re-keyed onto the surviving name when the target holds none; otherwise it would become unreachable and a fresh registered row would be minted in its place. When both hold one the target's governs and the dropped row's tier and seat counts are returned. 400 merging an org into itself, a confirm_org_key mismatch, or target_org_key_too_long_for:<table>.<column>; 404 source or target organization not found, including a non-UUID id (also when the flag is off); 412 migration_required:127-tenancy-fk-deferrable-for-merge.sql — a precondition, not a conflict; 503 the source org_key could not be resolved, so nothing was changed; 409 merge_conflict_unhandled:<table>. audit_log is deliberately NOT moved (audit rows are immutable, so the deleted org keeps its history; the count is returned as audit_rows_left_on_source), while users.primary_tenant_id IS repointed — it holds a tenant UUID under a non-standard column name and would otherwise dangle on the deleted tenant, silently demoting every merged user to their legacy per-user tier. Tier grants no access to this route — it is global-admin only, so an Enterprise subscriber who is not an admin is refused exactly as a Free one is. |
POST /internal/v1/admin/organizations | global-admin | Create org. Body: { slug, display_name, plan_tier, seat_limit }. 201. |
GET /internal/v1/admin/organizations/{org_id} | global-admin or org-manager | Get org. 404 when not found or out of scope. |
PATCH /internal/v1/admin/organizations/{org_id} | global-admin | Update org metadata. Only display_name and status are applied — plan_tier/seat_limit are rejected with 400 (use the licence endpoint), and status must be active | suspended | deleted (400 otherwise). Both fields are persisted to scf_tenants and tenant_admin_metadata in one primary transaction, with the response read back inside it, so a rename cannot half-apply and cannot return a stale replica read. |
POST /internal/v1/admin/organizations/{org_id}/license | global-admin | Set org plan tier + seat limit. Audited. Body: { tier, seat_limit, product }. product is vuln_intel | mpi; with enable_global_intel_license on it also accepts global_intel (tier restricted to pro|enterprise, 422 otherwise; plan_tier not mirrored). Flag off: 400 Invalid product. Changed 2026-10-01 (supersedes the 2026-08-17 seat floor), product: "global_intel" only: seat_limit is the number of bundle licence keys (one licence = one live phx_gintel_ key) and is required — omitted gives 422 seat_limit is required for global_intel. No tier default, no floor: 1 stays 1 and the response returns the stored value. 0 removes the licence (deletes the entitlement row, returns { "entitlement": null, "seat_limit": 0 }; issued keys keep working until revoked; no new keys can be minted). A value (including 0) below the sum of team reservations gives 422 Organization licences ({n}) cannot be below the team total ({m}). Lowering below the keys in use is allowed: existing keys keep working, new keys are blocked until used < number. vuln_intel/mpi keep optional seat_limit (omitted or 0 = tier default) and are unchanged by this endpoint. Added 2026-10-03 (migration 134-bundle-licence-expiry.sql), global_intel only: a licence can be a trial with an end date; when it ends the org has no licence and its bundle keys stop working (keys minted under a trial expire at the trial end at the latest). Optional body field licence_kind: "trial" starts the org's one 15-day trial with seat_limit (at least 1) licences (409 This organization already used its trial.; 409 This organization already holds a permanent licence.); "permanent" clears the end date with the tier and seat_limit sent (at least 1 licence; keys capped at a running trial get 90 days from the conversion; after the trial ended, expired keys stay expired; 404 Organization has no bundle licences. without a licence); omitted = the number is written and an existing end date is kept (on a trial that has ended, a positive number gives 409 The trial has ended. Use Make permanent to restore the licence., re-checked under the per-org licence lock on the database clock). seat_limit: 0 on an org that used its trial ends the licence instead of deleting it, so the one-trial rule holds. A licence_kind response adds licence: { tenant_id, tier, seats, kind, expires_at, trial_started_at, grant_source, active, used }. On an Orange-driven org (grant_source = 'orange', started through the PAI org grants) a global admin may change only the number: licence_kind, seat_limit: 0 or another tier gives 409 Licence is managed by Orange. Only the number can change. (re-checked under the per-org licence lock). The org listing adds global_intel_expires_at and global_intel_grant_source. Added 2026-08-26: for product: "vuln_intel" the entitlement and the org plan/seat mirror commit in one transaction, and the org's canonical identity is resolved before anything is written, so a refusal leaves no partially-applied licence — 404 no such org; 409 the org's normalised identity is already owned by another organization (scf_tenants.org_key is raw text, tenant_admin_metadata.slug is the normalised key, so keys differing only by case, spacing or confusables collide; resolve via the merge/repair workflow, no suffixed variant is minted); 422 the org key cannot be normalised into a valid identity (empty, over-long, or reserved such as admin/phoenix/default). |
| Endpoint | Auth | Purpose |
|---|---|---|
GET /internal/v1/admin/organizations/{org_id}/teams | global-admin or org-manager | List teams. Non-global admins see only teams in their scope. |
POST /internal/v1/admin/organizations/{org_id}/teams | global-admin | Create team. Body: { slug, display_name, seat_cap? }. 201. 409 SLUG_TAKEN on duplicate. |
GET /internal/v1/admin/organizations/{org_id}/teams/{team_id} | global-admin or org-manager | Get team. Caller must have the team in scope. |
POST /internal/v1/admin/organizations/{org_id}/teams/{team_id}/license | global-admin | Set team tier override + seat cap. Tier clamped to org tier. Body (every field optional for vuln_intel; seat_cap required for global_intel): { tier_override, seat_cap, product }. Added 2026-10-01: product is vuln_intel (default, unchanged) or global_intel; anything else is 400 Invalid product. For both products an unknown team, or a team of another org, is 404 Team not found in this org before any override is written. For global_intel (flag enable_global_intel_license; 400 Invalid product when off) the team reserves seat_cap bundle licences from the org number, tier_override is ignored, and seat_cap is required (422 seat_cap is required for global_intel). 422 Organization has no bundle licences. when the org number is 0/absent; 422 Team licences would total {n}, above the organization's {m}. when reservations would exceed the org number. The cap check and write run in one transaction under the per-org lock the key mint uses, so concurrent changes cannot over-allocate. Response: { "override": {...} }. |
DELETE /internal/v1/admin/organizations/{org_id}/teams/{team_id} | global-admin | Delete team. 204. |
GET /internal/v1/admin/organizations/{org_id}/membersAuth: Global admin or tenant_admin of that org. team_admin callers are auto-scoped to their own teams.
Query param: team_id (optional uuid) — filter by team. Overridden by auto-scoping for team_admin callers.
Response 200: { "members": [...] }
Each member carries the membership row plus the identity resolved from users
(2026-08-26):
| Field | Type | Notes |
|---|---|---|
user_sub | string | Cognito/basic-auth subject — the membership key. |
account_id | integer | null | users.id, joined on cognito_sub. Prefer this over user_id. |
user_email | string | null | Account email. This is the login identity — the platform stores no separate username for an established user. |
user_name | string | null | users.name; frequently blank in existing data. |
user_id | integer | null | Legacy. The scf_tenant_members.user_id column, which is not reliably populated — NULL on rows whose users.id exists. Kept for compatibility only. |
role | string | member | viewer | team_admin | tenant_admin. |
team_id, team_name | uuid | string | null | The member's single resolved team. |
joined_at | timestamp | Membership creation time. |
The identity join is a LEFT JOIN: a membership row whose sub has no users row is
still listed, with the three identity fields null, rather than disappearing from
its own org.
POST /internal/v1/admin/organizations/{org_id}/membersAuth: Global admin or tenant_admin of that org.
Add an existing user (by user_sub) to an org.
| Field | Type | Required | Notes |
|---|---|---|---|
user_sub | str | yes | Cognito user sub |
user_id | int | no | Internal DB user ID |
role | str | no | Default member. Values: member, viewer, team_admin, tenant_admin |
team_id | uuid | no | Assign to a specific team |
Errors: 400 invalid role; 403 TENANT_ADMIN_ESCALATION (tenant_admin caller cannot assign tenant_admin role); 409 seat limit.
Response 201: { "added": bool, "already_member": bool }
PATCH /internal/v1/admin/organizations/{org_id}/members/{user_sub}Auth: Global admin or tenant_admin of that org.
Update a member's role. Request body: { "role": "member|viewer|team_admin|tenant_admin" }
Errors: 400 invalid role; 403 TENANT_ADMIN_ESCALATION; 404 member not found; 409 LAST_ADMIN (cannot demote the last tenant_admin).
DELETE /internal/v1/admin/organizations/{org_id}/members/{user_sub}Auth: Global admin or tenant_admin of that org.
Remove a member. 409 LAST_ADMIN when removing the last tenant_admin. 204 on success.
PUT /internal/v1/admin/organizations/{org_id}/members/{user_sub}/teamReassign a member to a team. Enterprise orgs only (403 ENTERPRISE_REQUIRED for non-enterprise).
Request body: { "team_id": uuid }
| Endpoint | Auth | Purpose |
|---|---|---|
GET /internal/v1/admin/organizations/{org_id}/invitations | global-admin or org-manager | List invitations. team_admin auto-scoped. Query param: team_id. |
POST /internal/v1/admin/organizations/{org_id}/invitations | global-admin or tenant_admin | Create invite. Org admins cannot create tenant_admin invites. |
DELETE /internal/v1/admin/organizations/{org_id}/invitations/{invite_id} | global-admin or tenant_admin | Revoke invite. 204. 404 when not found. |
POST /internal/v1/admin/organizations/{org_id}/users — Org-scoped user creationCreate a brand-new user and enroll them in the org. Provisions the user in Cognito (sends invite email) and inserts a users DB row.
Auth: Global admin or tenant_admin of that org. team_admin callers always receive 403 TEAM_ADMIN_DIRECT_CREATE_FORBIDDEN.
| Field | Type | Required | Notes |
|---|---|---|---|
email | str (email) | yes | New user's email address |
full_name | str | yes | Display name |
role | str | yes | member, viewer, or team_admin. tenant_admin not permitted here. |
team_id | uuid | no | Assign to a team immediately |
industry | str | no | Default generic |
password | str | no | Temporary password; Cognito email invite sent regardless |
{ "user_sub": "cognito-sub-uuid", "role": "member", "org_id": "...", "team_id": null }
| Status | Code | Meaning |
|---|---|---|
| 400 | INVALID_ROLE | role is not member, viewer, or team_admin |
| 400 | — | Password does not meet Cognito policy |
| 403 | TEAM_ADMIN_DIRECT_CREATE_FORBIDDEN | Caller is a team_admin |
| 403 | TENANT_ADMIN_ESCALATION | Caller is tenant_admin and attempted to assign tenant_admin role |
| 404 | — | org_id not found or out of caller scope |
| 409 | SEAT_LIMIT | Organization has no available seats |
| 409 | USER_ALREADY_EXISTS | Email already registered in Cognito |
| Endpoint | Purpose |
|---|---|
GET /internal/v1/admin/org-scope | Return caller's org-admin scope. Response: { is_global_admin, is_org_manager, tenant_id, member_role, plan_tier }. Used by UI to gate the Organization-Team tab. |
GET /internal/v1/admin/member-teams | Map user_sub -> {org, team} across all tenants. Global admin only. |
Auth: require_admin (admin or global-admin Cognito group).
These are /api/v1/admin/ endpoints (not PAI). Documented here because their scoping behaviour changed as part of the org/team tenancy feature.
GET /api/v1/admin/registrationsList user registrations.
Scope: Global admin sees all rows. Non-global admins see only rows belonging to their primary org (get_user_primary_org). Non-global admins with no org membership receive an empty list.
Query parameters: status (optional), limit (default 100, max 1000), offset (default 0).
Response 200: Array of UserRegistrationListItem.
GET /api/v1/admin/registrations/pendingConvenience alias — equivalent to GET /registrations?status=pending. Same org-scoping rules.
Response 200: Array of UserRegistrationListItem with status=pending.
POST /api/v1/admin/registrations/{registration_id}/approveApprove a pending registration.
Auth: require_global_admin. Unlike the GET endpoints above this is not org-scoped — only global admins may approve.
Query parameters: notes (optional str) — free-text note recorded with the approval.
Response 200: MessageResponse — { "success": true, "message": "..." }. The message is composed at runtime and names what actually happened, including the organization the user was bound to and any binding warning. Treat it as human-readable prose, not a parseable contract.
Errors:
404 — no registration with that id.400 — registration is not pending (Cannot approve registration with status: <status>).502 — Cognito account provisioning failed. The registration is left pending and the call is safe to retry.Side effects:
approved and persists it via save_user_registration, which also writes the tenancy columns (tenant_id, requested_team_id, requested_role, intent) added by migration 102.users row, and clears the stored password.enable_approval_org_binding, default off): creates scf_tenants + a default team + tenant_admin/team_admin membership in one transaction and sets users.primary_tenant_id, so the org licence resolves. Every approved registrant gets their own org; a slug clash creates a numbered variant rather than joining an existing tenant. Binding is best-effort — a failure is reported in the response message and recorded on the registration, but never fails the approval. With the flag off, no tenancy rows are written.scf_seat_licensing_enabled is on. The licence key is the bound tenant's canonical org_key when binding succeeded and resolved one. When binding did not run or did not bind (flag off, no organization on the registration, or a binding failure), the key falls back to normalize_org_slug(registration.organization) — the same normalized form the tenancy tables use, never the raw form entry. Licensing is skipped entirely in two cases: the organization has no canonical slug (empty or reserved), or binding succeeded but the bound tenant's org_key could not be resolved — deriving one from the form entry there would licence a different key than the tenant actually uses.log_admin_action, and queues background notifications.docs/plans/2026-09-23-blue-intel-orange-endpoints-design.md. Feature doc:
docs/Individual_Feature/2026-09-23-blue-intel-orange-endpoints.md. Each endpoint below is
gated by its own SITE_CONFIG flag (default false, absent from site_config.json),
returns 404 when off, and requires a valid PAI key
(require_pai_or_global_intel for Endpoint A, require_pai_access for B/C/D) —
no new admission path is introduced.
POST /internal/v1/intel/library_versions/expanded — batch expanded library intelligence
Batch form of GET /internal/v1/intel/library_version/{purl}?include=... over up to 100 purls
in one call. Flag enable_intel_batch_expanded (also requires
enable_intel_query_surface; either off returns 404).
{ "purls": ["pkg:npm/lodash@4.17.20", "pkg:pypi/requests"],
"include": ["vulnerabilities", "malware", "exploitation", "campaign", "licensing"] }
library_version; a purl
without a version resolves as library (cross-version rollup).include must name at least one licence-gated ("expanded") token —
otherwise 422 include must name at least one expanded block.max_length).Response 200:
{
"items": [
{ "purl": "pkg:npm/lodash@4.17.20", "entity": "library_version", "status": "ok",
"envelope": { "...": "same envelope as GET .../library_version/{purl}" }, "block_errors": {} },
{ "purl": "pkg:bad", "entity": null, "status": "invalid", "envelope": null, "block_errors": {} }
],
"summary": { "requested": 2, "ok": 1, "not_found": 0, "invalid": 1, "timeout": 0, "error": 0, "units_charged": 1 },
"intel_license": { "...": "same stamp as the single-purl route, when licensed" }
}
status | Meaning | Charged? |
|---|---|---|
ok | Envelope built; some blocks may still be listed in block_errors. | yes |
not_found | resolve_entity found nothing for that purl. | yes |
invalid | parse_purl rejected the string. | no |
timeout | Per-purl deadline passed after starting, or the request deadline (intel_batch_request_deadline_ms, default 25000ms) passed before it started. | yes if started, no if not |
error | Unexpected exception scoped to that purl only. | no |
ok item still
carries its BASE envelope (no expanded block); no quota used, no expanded usage recorded.allow_clean is never passed — same invariant as the single-purl route.units_charged = purls that
finished ok, not_found, or timeout-after-starting.intel_batch_pai_max_abandoned_threads, default 3 — equal to the PAI
pool size, so a higher configured value is unreachable), a new batch request is refused before any work
and before any quota check:
503 { "detail": "batch capacity exhausted", "code": "batch_capacity" }
with Retry-After: 5.MAX_BATCH_PURLS = 100 is a hard
per-request cap; do not assume it can simply be raised to 500 from the caller's side. All purls share one
deadline (intel_batch_request_deadline_ms, default 25000 ms) and are composed by the PAI
lane's 3 workers, so the per-item budget is roughly deadline / (items / workers): about 0.75 s/item at 100
purls (25 s x 3 / 100) versus about 0.15 s/item at 500 purls (25 s x 3 / 500). A purl not yet started when
the deadline passes returns timeout (uncharged), so larger batches risk partial
timeout results once the licence, malware, vulnerability and PSS joins are all requested.
timeout items in a later
call.batch.check_capacity admits a request only while in-flight admitted requests plus
still-running abandoned threads on the PAI lane stay below
intel_batch_pai_max_abandoned_threads (default 3, per backend process), so a 4th concurrent
call, or fewer if earlier calls left abandoned threads, gets
503 {"code": "batch_capacity"} with Retry-After: 5. Treat that 503 as
retryable.intel_batch_pai_max_abandoned_threads and/or intel_batch_request_deadline_ms
for this route, measured against a real per-item latency benchmark (none exists in the repo yet). Do
not raise the cap alone.POST /api/v1/intel/library_versions/expanded — see
PUBLIC_API.html for the phx_gintel_/session-authenticated path.x-pai-token pair, not a session JWT, until ASPMAIN-6581 lands.POST /internal/v1/cves/batch — opt-in include intel blocks (added 2026-09-25)
Extends the existing batch CVE lookup mirror with an optional include list.
No include, or an empty list, is byte-identical to the pre-include
response — a hard invariant with its own golden test.
{ "cve_ids": ["CVE-2024-3094"], "include": ["ps_hp", "risk_band", "exploitation", "popularity", "github_pocs", "cwe_intel"] }
Gated by enable_pai_cve_batch_enrichment (default false). Off + include
set → 404; no include is unaffected. Unknown token → 400.
When include is non-empty, cve_ids is capped at 1000 →
422 above that; without include the existing flat 5000-id cap stays.
Each included block lands under a new top-level key intel on each item:
items[i].intel.ps_hp, items[i].intel.exploitation, etc. A failed block also carries
items[i].intel_errors.<block>: "source_unavailable" when the whole source is
missing, "error" when only that CVE's build raised. Source age is reported once, at the top
level, as intel_freshness.
ps_hp — source data/internal/high_profile_intelligence_internal.json.
"ps_hp": {
"status": "high_profile | watchlist | not_high_profile",
"ps_hp_score": 87.5, "ps_hp_tier": 1, "ps_hp_tier_name": "...",
"components": { "evidence": 0, "likelihood": 0, "severity": 0, "blast_radius": 0, "popularity": 0,
"reference_confidence": 0, "bugbounty": 0, "eol": 0 },
"hp_reasons": [], "hp_summary": "...", "hp_rationale": "...",
"enterprise_category": "...", "is_enterprise_watchlist": false
}
A CVE in enterprise_watchlist is watchlist; otherwise a CVE in
high_profile is high_profile. For a CVE not in the file, status
is "not_high_profile" and every other field is null — Blue never runs a
per-CVE fallback score for this block. Coverage today: 368 high-profile + 14 watchlist
CVEs. Treat not_high_profile as "no PS-HP data", never as "low risk".
risk_band — one band per CVE with its derivation named:
"risk_band": { "band": "CRITICAL|HIGH|MEDIUM|LOW|null", "derived_from": "hp_threat_type | cwe | none",
"threat_type": "...", "threat_impact": "...", "cwe_id": "CWE-79" }
Rule order: (1) PS-HP file's threat_type_risk when present; (2) otherwise the highest-risk CWE
from the CVE's own cwes list via web/data/cwe_catalog.json; (3) no source →
band: null, derived_from: "none".
exploitation
"exploitation": {
"kev": { "cisa": { "listed": true, "date_added": "2024-03-29", "due_date": "...", "ransomware_use": "Known" },
"vulncheck": { "listed": true, "date_added": "2024-03-29" } },
"epss": { "score": 0.94, "percentile": 0.99 },
"tooling": { "exploitdb_ids": [], "metasploit_modules": [], "nuclei_templates": [] },
"in_the_wild": { "verified_exploited": true, "exploited_sources": [], "first_exploited_ts": "..." },
"poc": { "first_poc_ts": "...", "has_public_exploit": true },
"exploitation_tier": "THEORETICAL | EXPLOIT_AVAILABLE | ACTIVELY_EXPLOITED | WEAPONIZED",
"tier_inputs_missing": ["zeroday"],
"predicted_exploitation": "POTENTIALLY_WEAPONIZED | POTENTIALLY_EXPLOITED | NONE | UNKNOWN"
}
tier_inputs_missing always lists "zeroday": no per-CVE zero-day source is
loaded, so exploitation_tier can never reach a verdict that needs that signal.
predicted_exploitation is EPSS-derived and is a separate axis — EPSS is never an input to
the tier.
popularity
"popularity": { "bug_bounty": { "listed": true, "total_reports": 12, "best_rank": 3, "popularity_score": 0.8,
"appearances": 4, "is_popular": true },
"news": { "hot": true, "mention_count": 40, "unique_sources": 9, "score": 0.7,
"first_seen": "2026-09-20", "window_date": "2026-09-22",
"top_articles": [ { "title": "...", "url": "...", "source": "...", "published": "..." } ] } }
At most 5 top_articles. Not listed/mentioned → {"listed": false} /
{"hot": false} with other fields null.
github_pocs — summary only (no repo list; use
GET /internal/v1/cve/{cve_id}/github-pocs for the detail drawer):
"github_pocs": { "poc_count": 7, "poc_level": "...", "has_github_poc_high": true,
"popularity_score": 0.6, "total_stars": 900, "total_forks": 120, "is_aggregate": false }
cwe_intel
"cwe_intel": [ { "cwe_id": "CWE-79", "name": "...", "threat_class": "...", "impact": "...", "risk": "HIGH" } ]
Unknown CWE → item with name/threat_class/impact/risk
all null.
Daily cap: each PAI key may enrich at most PAI_CVE_ENRICHMENT_DAILY_CAP
(default 500000) CVE ids per UTC day, counted by unique cve_ids on the request,
keyed by key_id in the shared cache (48h TTL). Over the cap:
429 { "detail": { "detail": "CVE enrichment daily cap reached", "code": "cve_enrichment_daily_cap" } }
(the detail field is itself a dict). Cache unavailable → the check fails
open and logs a WARNING.
Headers: X-Phx-Key-Kind: pai, X-Phx-Intel-Tier: none, plus, only
when include was set, X-Phx-Usage: cve_enrichment=<used>/<cap>.
POST /internal/v1/products/health — Product Health Score + fix evidence (added 2026-09-25)
Flag enable_pai_product_health (default false). Looks up PS-PHS/PS-PVS + NVD fix
evidence for up to 500 {vendor, product} pairs, from a new all-products artifact
(data/internal/product_health_index.json) built alongside — and without truncating
— the existing top-1000 product_vendor_intelligence_v2.json, which is unaffected.
{"products": [{"vendor": "Microsoft", "product": "Windows 10"}, {"vendor": "acme", "product": "gizmo"}]}
1-500 items, each field 1-200 chars. Normalisation before lookup: trim, lowercase, spaces →
_. Response 200, one item per request item, in the same order:
{
"items": [
{ "vendor": "Microsoft", "product": "Windows 10", "status": "ok",
"match": { "matched_by": "vendor_product", "key": "microsoft:windows_10" },
"health": { "phs": 41.2, "prs": 58.8, "grade": "D", "domain_scores": {}, "buckets": {}, "inputs": { "...": "the stored counts" } },
"fix": { "fix_available_rate": 0.62, "unfixed_count": 310, "unfixed_kev_count": 4, "evidence": "nvd_reference_tags" },
"watchlist": { "is_enterprise_watchlist": false, "cve_ids": [] } },
{ "vendor": "acme", "product": "gizmo", "status": "not_found",
"match": null, "health": null, "fix": null, "watchlist": null }
],
"provenance": { "generated_at": "...", "nvd_generated_at": "...", "source": "product_health_index" }
}
status: "not_found" with every other field null —
never omitted, never a fallback/zero score.health = the same calculate_product_score() formula
POST /internal/v1/calculate-phs-score uses.fix_available_rate = fixed_evidence_count / cve_count, rounded to 4 places; if
cve_count is 0 the whole fix block is null.{"detail": "product health index unavailable"} — never a 200 full of
not_found.X-Phx-Key-Kind, X-Phx-Intel-Tier).premium
product_health bulk export dataset,
gated by this same enable_pai_product_health flag. Both paths call the same scoring and
watchlist functions (pipeline/scoring/phs_scoring.py,
pipeline/product_health_lookup.py).POST /internal/v1/products/resolve-vendor — resolve NVD vendor for product names (added 2026-09-25)
Flag enable_pai_product_resolve_vendor (default false). Resolves the real NVD
vendor for a bare product name, sharing product_health_index.json and its 503/404 behaviour
with products/health above.
{"items": [{"product": "openssh", "vendor_hint": "multi-vendor", "cve_ids": ["CVE-2024-6387"]}]}
Rules, first match wins:
product empty after normalisation, or none/null → invalid.vendor_hint in {"", "none", "null", "multi-vendor", "multi_vendor"} counts as absent.hint:product in the index → ok, matched_by: "vendor_hint".ok, matched_by: "unique_product".cve_ids given: narrow via each CVE's own NVD CPE matches; exactly one left → ok, matched_by: "cve_ids".ambiguous, candidates ranked by cve_count.not_found.{ "items": [{ "product": "openssh", "vendor_hint": "multi-vendor", "status": "ok",
"vendor": "openbsd", "resolved_product": "openssh", "key": "openbsd:openssh",
"cpe_prefix": "cpe:2.3:a:openbsd:openssh", "matched_by": "cve_ids", "candidates": [] }],
"provenance": { "generated_at": "...", "source": "product_health_index" } }
The resolved vendor + resolved_product are valid input to
products/health as-is. Index missing → 503
{"detail": "product health index unavailable"}.
Unambiguous subset in bulk (added 2026-10-01, ASPMAIN-7801): the snapshot-only, base
product_vendor bulk export dataset
pre-resolves every product with exactly one vendor (rule 4). Ambiguous products are NOT in that dataset
and MUST be resolved via this route.
POST /internal/v1/usage/report — Orange's daily usage report (added 2026-09-25)
Flag enable_pai_usage_report (default false). Visibility-only: Orange reports its
own per-tenant daily query counts; quota balancing is explicitly not a goal.
{ "usage_date": "2026-09-22",
"rows": [ { "tenant_ref": "t_8f3a", "source": "PAI", "query_count": 1200 } ],
"seats": [ { "tenant_ref": "t_8f3a", "licence": "pro", "tier": "paid", "count": 5 },
{ "tenant_ref": "t_8f3a", "licence": "pro", "tier": "trial", "count": 1 } ] }
rows: 0-5000 entries (optional). At least one of rows or seats must be non-empty, otherwise 422; a body with only rows behaves exactly as before. tenant_ref matches
^[A-Za-z0-9._:-]{1,128}$. source in PAI, CONNECTOR,
CACHE. query_count 0-1,000,000,000. usage_date not in the future
and not older than 35 days. A repeated (tenant_ref, source) pair → 422.seats (optional, added 2026-10-03): 0-5000 entries, one per
(tenant_ref, licence, tier). tenant_ref has the same rule as in rows.
licence and tier are lowercase tokens matching
^[a-z0-9][a-z0-9._-]{0,63}$, an open set: Blue stores them and does not
interpret them, so a temporary tier: "trial" seat is accepted. count is an integer
0-1,000,000 (a day with seats but no usage rows is valid). A repeated
(tenant_ref, licence, tier) triple → 422.pai_consumer_seat_reports (migration
133-pai-consumer-seat-reports.sql), unique on
(key_id, usage_date, tenant_ref, licence, tier); replace-on-conflict, so a retry never
double counts.pai_consumer_usage_reports (migration
129-pai-consumer-usage-reports.sql), unique on
(key_id, usage_date, tenant_ref, source).rows and seats is written in ONE
transaction. If either write fails (for example before migration 133 is applied), neither is stored and the
request fails, so a retry is safe.200 {"usage_date": "...", "accepted": <n rows>, "seats_accepted": <n seats>} (each is 0 when that section was not sent). 404 when the flag is off; 503 when the storage method for the sent sections is unavailable (nothing is written)./internal/v1/grants/*, added 2026-10-03)
Orange starts a Global Intel trial for one of its tenants; Blue creates (or finds) one organization
for that tenant, gives it a temporary global_intel licence and, on convert, makes the same
organization permanent. When a trial ends without a convert, access ends: the licence row carries the end date and
the organization's Intel Bundle keys are capped to it.
Gating: flag enable_pai_org_grants (default false) and
enable_global_intel_license. Unless both are on, every route below returns 404 (not 403);
the flag check runs before authentication. Auth: PAI key with scope pai_internal (plus the
usual PAI IP allowlist / companion token). Per-key scoping: a key only sees the grants it created; a grant
created by another key and an unknown grant_id are the same 404 not_found.
| Method | Path | Purpose |
|---|---|---|
POST | /internal/v1/grants/trial | Start a trial (creates the organization on first use) |
POST | /internal/v1/grants/{grant_id}/convert | Make the grant permanent (same organization; also after expiry) |
GET | /internal/v1/grants/{grant_id} | Read a grant (an expired trial is 200 with state: "expired") |
Start body:
{ "tenant_ref": "orange:t-8f3a", "tier": "pro", "seat_limit": 1,
"expires_in_days": 15, "idempotency_key": "start-8f3a-0001" }
tenant_ref matches ^[A-Za-z0-9._:-]{1,128}$ (opaque, never a name). tier is
pro or enterprise. seat_limit optional integer 1-1000, default 1
(the number of Intel Bundle licence keys the organization may mint). expires_in_days optional integer
1-30, default 15. idempotency_key 8-128 chars of [A-Za-z0-9._:-]. Unknown
fields → 422.idempotency_key, same body).(key, tenant_ref), ever: a second start with a new idempotency_key →
409 trial_exists with the existing grant_id and its state,
whether the trial is active or expired.PAI_GRANTS_DAILY_START_CAP (default 50) starts; over it →
429 daily_start_cap.PAI_GRANTS_MAX_ORGS_PER_KEY (default 1000) organizations ever; a start that would
create one more → 429 org_ceiling_reached (a start for a tenant_ref that
already has its organization is not counted). Both limits are checked before the organization is created.tenant_ref is case sensitive: Foo and foo are two different tenants.org_key pai-<16 hex>, or
pai-<16 hex>-<8 hex> after a name collision, display name
PAI trial <tenant_ref>), a default team, and no users, no members and no keys.
Convert body: { "tier": "pro" | "enterprise", "seat_limit": 1-1000 (optional, omitted keeps the
current value), "idempotency_key": "..." }. 200 with kind: "permanent",
state: "permanent", expires_at: null, same organization_id. Works while the trial
is active and after it expired (reactivates; the organization, its users and data are untouched).
Converting an already permanent grant again with new values updates them; the same values are a no-op
200. Per key per UTC day, at most PAI_GRANTS_DAILY_CONVERT_CAP (default 100) converts; over
it → 429 daily_convert_cap. An idempotent replay of a convert already recorded is
always 200 and never counts.
Not a request gate: a grant is the licence record for the organization. The organization has no
members and no keys and the consumer calls Blue with its own PAI key, so Blue does not refuse a
consumer's PAI calls for a tenant_ref when its trial ends. A consumer that must stop serving an expired
trial gates on its own side using expires_at and state
(GET /internal/v1/grants/{grant_id}).
Response (all three routes):
{ "grant_id": "7c1e...", "tenant_ref": "orange:t-8f3a", "organization_id": "0b6d...",
"product": "global_intel", "kind": "trial", "state": "active",
"tier": "pro", "seat_limit": 1,
"starts_at": "2026-10-03T12:00:00Z", "expires_at": "2026-10-18T12:00:00Z" }
state: permanent (kind permanent), active (trial, licence still live),
expired (trial ended, or the licence was removed by an operator; then tier/seat_limit
are null). tier, seat_limit, kind, starts_at,
expires_at and state are read from the organization's licence, which is the authority;
timestamps are ISO 8601 UTC.
Idempotency: (key, idempotency_key) is unique across start and convert and stored with a
SHA-256 hash of the request. Replaying the same request returns 200 with the grant's current
state; the same key with a different body (or on a different route/grant) → 409
idempotency_conflict. A start that crashed half way is completed by its retry with the same
idempotency_key.
Errors (body {"detail": "...", "code": "..."}; trial_exists adds
grant_id and state):
| Status | code | When |
|---|---|---|
| 404 | (none, {"detail": "Not Found"}) | either flag off |
| 404 | not_found | unknown grant, another key's grant, malformed grant_id, licence reports no_licence / unknown_organization |
| 409 | trial_exists | this (key, tenant_ref) already had a trial, including one a Blue admin ended early or one a Blue admin started on the organization (never adopted as this grant; body carries grant_id, state) |
| 409 | idempotency_conflict | idempotency_key reused with a different request |
| 409 | convert_superseded | a newer convert exists for this grant, so this older request was not applied; every retry of the same idempotency_key stays 409, so send a new request with a new idempotency_key |
| 409 | licence_already_active | the organization already holds a live permanent licence, so a trial is refused |
| 422 | (FastAPI validation body) | request body validation |
| 422 | invalid_tier / invalid_seats / invalid_days | refused by the licence service |
| 422 | org_below_team_caps | convert seat_limit below the organization's team caps total |
| 422 | other licence code | any other licence refusal (logged) |
| 429 | daily_start_cap | per-key daily start cap reached |
| 429 | daily_convert_cap | per-key daily convert cap reached (PAI_GRANTS_DAILY_CONVERT_CAP) |
| 429 | org_ceiling_reached | per-key lifetime organization ceiling reached (PAI_GRANTS_MAX_ORGS_PER_KEY) |
| 500 | internal_error | licence service answered invalid_source or licence_managed_by_orange (a Blue bug, logged) |
| 503 | storage_unavailable | storage or licence service unavailable |
Storage: pai_org_refs, pai_org_grants, pai_org_grant_converts (migration
135-pai-org-grants.sql); the licence itself is tenant_product_entitlements
(global_intel), written only through the bundle-licence service (migration
134-bundle-licence-expiry.sql). An organization created by a grant cannot be merged by a Blue admin
(refused with merge_blocked:pai_org).
Every route above sends response headers reporting the caller's key kind, licence tier and (where applicable) daily usage — never as body fields:
| Header | Value |
|---|---|
X-Phx-Key-Kind | pai or gintel (public twin only) |
X-Phx-Intel-Tier | pro, enterprise, or none |
X-Phx-Usage-Date | UTC date the counters belong to, YYYY-MM-DD |
X-Phx-Usage | intel_expanded_lookups=<used>/<limit>, intel_query_lookups=<used>/<limit> on Endpoint A; cve_enrichment=<used>/<limit> on Endpoint B |
Endpoint A (PAI + public), GET /internal/v1/intel/*, GET /api/v1/intel/* get the
full set. Endpoint B and products/health/products/resolve-vendor have no daily
counter of their own, so only the tier-only pair is sent. If the usage row cannot be read,
X-Phx-Usage is omitted entirely (never a fabricated 0).
X-Phx-Usage is served from a 30-second per-subject cache: this value may lag your most
recent calls by up to 30 seconds; it is provided for visibility only and never affects request
acceptance — quota is enforced separately and always against a live count.
GET /internal/v1/enterprise-watchlist — paginated list of the PS-HP
file's enterprise_watchlist table, keyed by CVE. Representative fields:
cve_id, cpe, enterprise_cpe, enterprise_vendor,
enterprise_category, enterprise_risk_score, kev_vendor,
kev_product, technology, cvss, epss,
ps_hp_score, hp_reasons, plus the GreyNoise/Shadowserver/TTE/EOL fields
high_profile carries. No purl field — this table is
CVE-keyed, not package-keyed.
POST /internal/v1/calculate-phs-score — PS-PHS is the
Product Health Score, PHS = 100 - PRS, with letter grades A-F. This route is a
calculator: it computes a score from counts the caller supplies, never a lookup by name.
For a lookup of a real vendor/product's PHS, use POST /internal/v1/products/health
above — it resolves the counts from Blue's own corpus and calls this same formula.
GET /internal/v1/phoenix-score/{cve_id} — full PS-HP field set for one
CVE. Coverage caveat: only 368 high-profile + 14 watchlist CVEs (382 total) are in the PS-HP
file — a CVE outside it falls through to a computed score built from
get_enriched_cve()'s KEV/EPSS/GitHub-PoC enrichment, a materially different code path from the
curated PS-HP fields.
Field-name pin: mitre.mitre_technique_ids. The canonical PAI field name (the
expanded envelope) is mitre.mitre_technique_ids. Existing public field names using
a different spelling for the same concept are not renamed — this doc is the single
place that states which surface uses which name.
PAI calculation endpoints apply neutral defaults when parameters are omitted:
null and excluded via renormalization.unknown status.