The catalog-wide bridge in EmexSourceDbService.fetchCategoryParts was
measured against vehicle_parts on 2026-06-01 and found to return 7-114x
more parts than belong to the requesting vehicle, with 49-98 wrong OEM
codes per 100 served. That directly violates the project rule that the
user must never see a wrong OEM.
Per-catalog noiseRatio sample (catalog-wide / per-vehicle):
RENAULT201910 51x | FFIAT84 45x | VOLVO201410 24x | MB201810 14x
AU1587 8x | BMW202501 70x (+ gid namespace mismatch ETK vs numeric)
GM_C201809 114x | MINI202501 12x | LRE201412 7x | MAZDA2020 54x
GM_OP201809 dump has only 1 wildcard vehicle (unique_key="_") so the
single Crossland X "owns" all 47k Opel parts — same firehose served
to any Opel sub-model in sase prod.
All alternative bridges were proven dead:
SSD eşleştirme - session-bound, 0/91 sase SSDs match dump
scrape_queue_v2.vehicle_ssd - same session SSD format
api_cache replay - table empty (0 rows)
wizard_parameters - table empty (0 rows)
VIN direct - no VIN column in dump
The only viable per-vehicle bridge is vehicles.unique_key reconstruction
from raw_data.parsedOptions, but sase currently stores the required 4
wizard fields on just 5/103 emex vehicles (all Renault). That work is
follow-up; this patch only stops the bleeding.
Change:
- Add EMEX_SOURCE_DB_ALLOWED_CATALOGS env (comma-separated, default "")
- EmexSourceDbService.fetchCategoryParts returns null unless catalogCode
is in the allowlist. Empty allowlist = service is effectively off for
parts, full fallthrough to live emex.
- Connection pool stays alive so the follow-up per-vehicle bridge /
schema-only path can use it without flipping env.
- Boot logs warn loudly when connected with an empty allowlist.
Prod was never affected — CATALOG_SOURCE_DB_ENABLED was unset there. This
fixes dev branch behaviour (default-on since commit 3a3a7d3) and keeps
prod safe by default once main is promoted.
Files:
- packages/config/src/index.ts env schema + audit notes
- apps/api/src/config/configuration.ts parse allowlist into string[]
- apps/api/src/integrations/catalog-source-db/emex-source-db.service.ts
allowlist field, init logging, fetchCategoryParts gate, class doc
- docker-compose.coolify.yml env injection for api + worker
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Verified 2026-06-01 against dev's 103 unique pcat carIds: the current pcat
dump's deep-scrape (7.978 cars with real parts data via schema_parts or
part_groups+part_group_items) targets a US/JDM-market subset — Toyota 2112,
Nissan 1508, Audi 1311, Chevy 1050, Hyundai 745. **None** of sase's TR-market
vehicles intersect that rich subset:
- 18/103 sase carIds are in dump.cars at all (registry only)
- 0/103 yield parts via Bridge A (schema_images → schema_parts)
- 0/103 yield parts via Bridge B (part_groups → part_group_items)
Even the cars that match by exact carId (Fiat Doblo 368 schemas, Renault
Megane, Bravo 456 schemas) have only diagram metadata — no parts annotation.
The dump scraper finished tier-1 (catalog/model/car listing) and tier-2
(schema diagrams) for these, but stopped before tier-3 (parts annotation).
Under the strict "always correct OEM" constraint there is no safe pcat lookup
today. Disable it. The container stays up for future use cases (OEM cross-
reference search, alt-part matching) and so we can flip the env back without
a code change if a richer dump arrives.
EMEX stays on (its catalog-allowlist is the next step).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Initial design routed emex lookups through vehicles.rawData.ssd → dump vehicles
→ vehicle_parts. Smoke test against prod ssd values: 0 / 10 matched. EMEX
regenerates the SSD on every decode session, so sase's stored SSD never
matches the SSD the dump scraper recorded for the same physical vehicle.
Pivot to a catalog-wide bridge that actually works:
catalogs.code ↔ vehicles.rawData.catalogCode (e.g. "RENAULT201910")
part_groups.group_id ↔ categories.externalId (e.g. "11754")
→ parts via vehicle_parts.group_id (dump's parts.group_id is 100% NULL)
Verified coverage on prod's 8287 unique (catalogCode, gid) pairs: 25/26
catalog codes resolve, 7178 pairs hit a part_group (87%), 5919 of those
return actual parts via vehicle_parts (~71% net). Tradeoff: returns all
parts in the (catalog, group) across every variant in the catalog, so the
result is slightly noisier than the live per-vehicle scrape. Acceptable —
parts overlap heavily and the upstream-call savings outweigh the noise.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds an optional local-dump lookup layer in front of the live PartsCatalogs and
EMEX scrapes. When enabled, getCategoryWithPartsInner queries a Postgres
(pcat) or MariaDB (emex) dump for the requested schema/group's parts and
hotspots; on miss it falls through to the existing upstream call unchanged.
Hits avoid the live API, its cooldown, and its rate-limits — direct DB latency.
- New CatalogSourceDbModule with PcatSourceDbService + EmexSourceDbService
(raw SQL, no Drizzle schema modeling — dump shapes are frozen snapshots).
- pcat lookup keys on schema_images.schema_ext_id (the dump's column that
matches sase's pcat groupId; observed ~7% hit rate on prod's 6596 unique
groupIds, of which ~10% have schema_parts → ~3-5% net parts coverage).
Joins schema_parts → parts directly; the dump's part_groups+part_group_items
linkage covers 0 of our hits, so we skip that path entirely.
- emex lookup uses (catalog_id, ssd) → vehicles.id then (vehicle_id, group_id)
→ vehicle_parts → parts + part_images. The ssd is already persisted into
vehicles.rawData.ssd by the existing emex.mapper, no extra capture needed.
Gated behind CATALOG_SOURCE_DB_ENABLED + PCAT_SOURCE_DB_URL / EMEX_SOURCE_DB_URL.
All three default unset, so this commit is a no-op until prod env is configured.
Adds mysql2 dep for the MariaDB client.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>