Commit Graph

263 Commits

Author SHA1 Message Date
26fae98246 fix(emex): coerce proxy port env to number (boot-crash on real range)
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
ConfigService.get<number>("EMEX_PROXY_PORT_START") returns the raw env
STRING; the port-pick arithmetic then string-concatenated it
(45 + "10001" = "4510001"), producing an out-of-range port that made
undici's `new URL` throw "Invalid URL" at EmexService construction —
crashing the entire API on boot.

A single-port range (823) happened to concat to a still-parseable "0823",
which masked the bug for months. It surfaced the moment the prod
EMEX_PROXY_PORT range was widened (823 -> 10001-10099) to let Q1's
per-request port rotation work: prod crash-looped until the env was
reverted. Coerce to a validated integer port (1-65535) with default
fallback so a real range is safe.

Regression test: constructing EmexService with string port env over a
real range must not throw.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 23:04:49 +03:00
a3ae90ec29 fix(vin-decode): 4 RCA-confirmed decode-chain bugs
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Root-cause analysis live re-decoded all 124 historically-undecoded prod
VINs; 38 already decode now. These 4 fixes target confirmed code bugs
that drop or mask real decodes (see undecoded-vin-rca.md):

Q2 — PL24 circuit breaker now only counts transient transport faults. A
definitive upstream negative (NotFound/BadRequest) no longer trips the
global 30s breaker that was starving PL24 for every subsequent VIN
(the sibling-VIN inconsistency in the report). Live-proven on VR7.

Q3 — previewVin / multi-candidate path no longer returns an empty
success: the pcat/emex candidate branches fill brandName (from catalogId
/ WMI), fixing the 6 "HTTP 200 with null brand+model" cases.

Q1 — EMEX fetch retries transient proxy failures with a FRESH ProxyAgent
per attempt (rotates the DataImpulse port; ~42% blip rate observed),
plus an opt-in direct fallback (EMEX_DIRECT_FALLBACK). HTTP answers are
never retried.

Q4 — VIN resolve cache keys namespaced by DECODE_CHAIN_VERSION and the
negative TTL drops 6h -> 30m, so a decode-chain fix self-heals stale
negatives on deploy instead of masking phantom-undecoded VINs for hours.
The admin cache-buster uses the same key builder.

Tests: 179 passed (+ new Q2/Q3/Q4 specs). typecheck + biome clean.

Deploy note: prod EMEX_PROXY_PORT_START/END are both 823 (single port);
widen to a real range (e.g. 10001-10099) in Coolify so Q1's port
rotation takes full effect.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 22:37:11 +03:00
bb60a95ec3 perf(backfill): Phase-1 residue exclusion — stop re-picking dead vehicles (Tier 3)
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Phase-1 of the backfill scan selects all zero-parts decoded vehicles every wave.
Genuine residue (VINs with no catalog data anywhere — model-indexed HKN,
EMEX-uncovered, etc.) stays zero-parts forever, so it filled the batch every
wave, re-attempting dead vehicles and starving the Phase-2 rolling rescan (its
cursor was stuck for a week).

Track a per-vehicle no-result counter (prefetch:noresult:<id>) incremented when a
prefetch attempt finishes with the vehicle still at zero parts (0 categories in
init, or 0 parts after the whole chain). tryPick skips vehicles past
PREFETCH_NORESULT_MAX (default 2) attempts; the counter has a TTL
(PREFETCH_NORESULT_TTL_DAYS, default 7) so a later catalog fix re-fills them.
Frees capacity for fillable vehicles and lets Phase-2 run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 19:15:57 +03:00
48aea65b5c perf(backfill): raise prefetch worker throughput, env-tunable (Tier 2)
After Tier 1 removed the self-throttle, the 5 jobs/min limiter + concurrency 1
became the bottleneck. Make throughput env-tunable so prod can ramp while
watching the fail rate:
- concurrency 1 -> 3 (PREFETCH_CONCURRENCY): parallelises emex/pl24 so a slow
  parts-catalogs job no longer head-of-line-blocks the queue.
- rate ceiling 5 -> 20 jobs/min (PREFETCH_RATE_MAX).
- parts-catalogs: drop the pathological cumulative index*20s enqueue delay (the
  Nth leaf of a vehicle waited N*20s); keep one per-job pace (PCAT_PACE_MS,
  default 15s, 0 to disable).
- PL24 09-18 scrape window now env-tunable (PREFETCH_PL24_START / _END).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 18:43:10 +03:00
1e3953417e perf(backfill): stop the prefetch worker from throttling itself (Tier 1)
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
The worker's own upstream fetches called touchActivity(), setting the
prefetch:activity:<source> cooldown key (TTL 300s) that checkCooldown then
honoured — so after each fetch the worker paused itself for up to ~5 minutes
(re-checking every 60s, ~5 empty cycles per key). At ~1 fetch / 5 min the
3387-job backlog needed ~6 days to drain.

- Wrap each worker job in an AsyncLocalStorage backfill context; touchActivity
  skips the cooldown key when invoked from the worker, so the cooldown reflects
  only real user requests (worker yields to users, never to itself).
- Cooldown TTL 300s -> 90s (a request 5 min ago isn't "active").
- checkCooldown pauses for the key's actual remaining TTL (one wait) instead of
  a fixed 60s re-check loop.

No extra upstream load — only removes the worker's self-imposed idle time.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 18:23:54 +03:00
3bfefe6196 feat(analytics): canonical subscription_activated revenue event ()
No PostHog event carried realized revenue, and EFT/havale activations fired
nothing at all — so total paid revenue / MRR was unmeasurable (a Stripe DWH
connector alone would also miss EFT). activateSubscription is the shared
chokepoint for both Stripe (stripe.service) and EFT/manual (billing.service)
activation, so emit one canonical subscription_activated there with PostHog
revenue props: $revenue (major TRY), currency, mrr (yearly amortised /12),
plan, plan_id, brand_count, billing_period, method (looked up from the latest
payment row), referral_credit_days. Funnel steps keep their kuruş 'amount' but
intentionally carry no $revenue, so revenue isn't double-counted.

Unblocks trial->paid, MRR/ARPU and revenue-by-plan/channel across ALL payment
methods. Injected PostHogService (PostHogModule is @Global).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 16:02:30 +03:00
c5f8288d71 fix(catalog): case-insensitive PL24 group-wid drill check
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
PL24 sub-group nav nodes were classified by linkWid.includes("Group")
(case-sensitive). That matched capitalised wids (subGroupsTable) but
missed lowercase ones — groupReferenceTable, groupTable, groupsTable
(~1157 leaf nodes in prod) — so those skipped the group-drill branch in
getCategoryWithPartsInner and the reference-resolution descent, falling
to the parts path (a wasted upstream fetch; the generic drill-on-empty
fallback then re-drilled them). Lowercasing the check routes these nav
nodes straight to children/loadError like their capitalised siblings.

Empty-catalog audit (2026-06-03) showed PL24 drives 61% of user-seen
'0 parça' views; pcat fake-leaves are effectively solved (1 case).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 15:40:39 +03:00
642139de9c fix(pl24): detect "bakınız tablo, konum:" reference phrasing too
PL24 translates "see table" two ways — "bk. tablo:" and "bakınız tablo,
konum:". Detection only matched the first, so the latter rows (e.g. evaporator
housing → 820-020) stayed dead. Broaden the name regex to match either, and
strip both phrasings from the displayed label. The code-in-remark gate still
prevents flagging real parts that merely mention "tablo".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 20:20:30 +03:00
c440252aa3 feat(pl24): on-demand drill to resolve unseeded "bk. tablo:" references
When a reference's target illustration isn't seeded yet (load-time index
miss → categoryId null), clicking it now calls a new resolve endpoint that
drills the relevant main-group root (its external_id = the code's first
digit; the illustration is a direct child) and re-resolves. One PL24 call in
the common case, bounded + cached; falls back to pre-filled search if not
found. UI shows a spinner on the button while drilling.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 19:41:02 +03:00
232c9ccc7a feat(pl24): make "bk. tablo:" BOM cross-references navigable
PL24 BOM emits "see table NNN-NNN" reference rows (oem N/A, target code in
remark) with NO upstream link. Resolve the code against the vehicle's
illustration index (codes live in category names as {NNN-NNN}) and render
jump links. Unresolved targets (branch not seeded yet) deep-link a pre-filled
catalog search via ?q=.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 18:55:42 +03:00
00f7edbc15 fix(pl24): Volvo vin-image-board BOM parts + decode subgroup name entities
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Volvo's vin-image-board.action ships its BOM inline as partno= tc-data-row
rows (no pncHierCode / json-vin-bom-detail), so the Ford VIN-BOM parser
returned 0 parts. Fall back to parsePsaBomParts (partno= rows) when no pnc
rows are found. Also decode HTML entities (&Ouml; &quot; …) in scraped Volvo
subgroup names. Completes the Volvo chain: group1→group2→illustration→parts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 02:30:54 +03:00
c46695d320 fix(pl24): Volvo drill — keep openVinDialog=false links + image-board leaves
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Two bugs in the vin-group.action subgroup scrape: (1) the filter dropped any
href containing "openVinDialog", but the real sub-group links carry
openVinDialog=false (only the VIN-dialog crumb is =true) — so every child was
discarded; (2) the deepest group level lists its illustration leaves as
vin-image-board.action links, which weren't extracted. Now match both deeper
vin-group.action?groupN= and vin-image-board.action anchors, and only drop the
openVinDialog=true crumb. Completes Volvo group1→group2→illustration→parts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 02:24:59 +03:00
f5c3198a6e fix(pl24): demo-page retry in fetchP4Page + Volvo drill diagnostics
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
PL24 serves a stripped NOT_LOGGED_IN_DEMO page (no groups/parts) when the
service token is stale. decodeVinForService retries on this, but drill paths
(fetchSubGroupsByPath/fetchPartsByPath) reach upstream only via fetchP4Page,
which didn't — so Volvo subgroup drilling parsed empty demo pages. Retry once
with fresh auth on a demo page. Adds a Volvo-drill diagnostic log.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 02:18:04 +03:00
7208069a80 fix(pl24): prefix basePath for relative P4 hrefs in fetchP4Page
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Volvo vin-group.action categories store hrefs relative to the catalog dir
(e.g. "vin-group.action?group1=…"). fetchP4Page did `${baseUrl}${url}`,
collapsing to "partslink24.comvin-group.action" → ENOTFOUND. Prefix the
service basePath when the path is relative. Fixes Volvo subgroup drilling
for existing (relative) stored linkPaths without a re-decode.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 02:12:59 +03:00
10dcf2d6ae fix(parts): dedupe ingest path + clean ~108k duplicate rows
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Backfill + reactive prefetch can re-drill the same category multiple
times, and the parts insert path had no dedupe guard. Result: 8.2%
duplicate rows on pl24, 14.5% on parts-catalogs, and 33.7% on emex —
~108k extra rows across 7,571 categories on 224 vehicles. Every drilled
catalog page rendered each part twice (the Tampon example: 32 rows for
19 distinct OEMs).

* Migration 0010 — phase 1 deletes existing dupes preserving the
  oldest row per (vehicle_id, category_id, oem_code, name, position)
  group; phase 2 adds a UNIQUE INDEX over the same tuple with NULLS NOT
  DISTINCT (PG 15+) so null position/vehicle_id collapse like equal
  values rather than each counting as its own "distinct" row. Idempotent
  CREATE UNIQUE INDEX IF NOT EXISTS so the runner is safe to re-apply.
* All five insert(parts).values(...).returning() call sites
  (parts.service, categories.service ×3, catalog.service) get
  .onConflictDoNothing() so future re-drills no-op instead of erroring
  on the new constraint. `.returning()` continues to surface only the
  newly-inserted rows; existing logs read `Stored N parts` as actual
  net insertions, which is what we want.

Dry-run on dev DB: 524,540 → 425,186 parts (99,354 dupes deleted), index
created cleanly. Same delta expected on prod (~108k drop).

drizzle-orm 0.41 doesn't expose .nullsNotDistinct() on the index builder
so the constraint is owned by raw SQL — see the inline comment in the
parts schema and the migration file. Future schema generators should NOT
try to drop or rewrite this index.
2026-06-02 02:08:57 +03:00
d792bf0efb fix(pl24): Volvo subgroup drilling via vin-group.action HTML
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Volvo's VIN catalog is a 3-level vin-group.action?group1=…[&group2=…] HTML
tree; PL24's json-vin-*-group.action JSON endpoints now 404. Add HTML
subgroup extraction (keep links one group-level deeper than the current
path) in ford-legacy fetchSubGroupsByPath, and route vin-group.action?group1=
nodes through getChildren (drill-first, fall back to parts) in
getCategoryWithParts. Recovers Volvo group1→group2 navigation. Leaf parts
handled separately.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 02:08:12 +03:00
316a5e014c feat(backfill): widen scrape window for pcat + double batch size
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Backfill was sweeping only 20 vehicles/hour and skipping pcat outside
09:00–19:00 Istanbul, leaving the catalog backlog (115 pl24 + 63 pcat +
27 emex zero-parts vehicles) crawling forward at ~528 parts/day and the
biggest user vehicle (Toyota Corolla 2026) unchanged across 24h.

Two related changes — both unblocked by the Redis-persisted warm JWT pool
(aa4d055 + dcf7e06) which keeps pcat captures alive 24/7 instead of
needing the 09:00–19:00 office-hours assumption:

* isWithinTimeWindow: parts-catalogs no longer gated — the warm pool +
  Redis hydration cover the cold-start case the old office-hours rule
  was working around. PL24 keeps its 09:00–18:00 window because the
  upstream rate-limit is still tighter outside it.
* BACKFILL_BATCH_SIZE 20 → 40 — twice as many vehicles per wave, still
  protected by MAX_BACKLOG=1000 self-throttle and per-source cooldown
  (prefetch:activity:<source>) so live-user traffic still gets priority.

Combined effect: pcat goes from ~10h/day to 24h/day, batch doubles —
roughly 3× backfill throughput. Worst-case proxy spend tracked by the
DataImpulse daily cap; if a wave saturates upstream the rate-limit
handler (c7e59b9) defers without burning attempts.
2026-06-02 02:00:16 +03:00
0f647b3931 fix(pl24): recover Volvo categories — stop vin-group.action?group1= shadowing
Volvo (VIN-indexed legacy catalog) ships its real top groups as
vin-group.action?group1=… in the vin-group HTML. PL24's
json-vin-main-group.action endpoint now 404s, so decode falls back to the
HTML scrape — but parseP4NavigationCategories AND the seed-time NAV_CRUMB
filter both blanket-exclude vin-group.action, dropping every real Volvo
group → 0 categories. Exclude vin-group.action as a crumb only when it
lacks group1=. Confirmed upstream: 7 real groups (Frenler, Elektrik
sistemi, …) present in the HTML for YV1AS7050A1118639.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 01:51:19 +03:00
ed74e2f361 feat(register): VIN-aware teaser + B2B copy
Hero already shows a generic vehicle preview when a 17-char VIN is typed,
so /register?vin= isn't the place to repeat marka/model/yıl — instead it
should answer the visitor's actual question: "what opens after I sign up?"

Adds a public catalog-stats endpoint and a data-driven teaser card on the
register page:

Backend:
* GET /api/vehicles/:vin/teaser-stats (Public, VIN-validated). Single SQL
  round-trip counts categories + parts + schema_pics for the VIN. Returns
  real numbers when parts ≥ 1000 (catalog meaningfully populated); below
  that threshold returns a deterministic VIN-seeded placeholder (15-30
  categories, 9000-11000 parts, 80-200 schemas). Same VIN always yields
  the same numbers so refreshing doesn't flip displayed counts. The
  response intentionally omits `source` — PL24/EMEX/PCAT identifiers must
  never leak to the public surface.

Frontend (/register):
* When ?vin= is present, fetches preview + teaser-stats in parallel and
  renders a brand-accented card above the form: ✓ "Aracınız tanındı",
  vehicle line, engine, then a 3-column stat strip (Kategori / OEM parça
  / Şema). Below: "Hesap açtığında bu araç için kataloğa anında erişim
  açılır."
* B2B copy pass on the rest of the page:
  - Heading flips to "Hesap Aç ve Katalogu Gör" when VIN present
  - Trial messaging rewritten to anti-gimmick B2B tone:
    "Kart bilgisi gerekmez · 30 gün ücretsiz · istediğin an iptal"
    (was: "30 gün Full Paket ücretsiz deneyin — kredi kartı gerekmez")
  - Subhead: "Sınırsız şase sorgulamak için ücretsiz hesap aç"
  - Submit button: "Hesap Aç ve Katalogu Gör" (vin) / "Hesap Aç" (no vin)
  - "Ücretsiz Başla" / "Full Paket" strings purged per [[sase-b2b-copy-not-consumer]]
2026-06-02 01:07:18 +03:00
5e9d9050b9 fix(pl24): stop .action shadowing PSA illustration/parts dispatch
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
isP4LegacyPath matched any ".action" path, shadowing the PSA
json-illustrations.action / image-board.action dispatch in
fetchSubGroupsByPath and fetchPartsByPath. Every PSA (Citroen/Peugeot/DS)
drill below main-group level fell through to the Ford/Fiat legacy parser,
which cannot parse PSA JSON, so it returned empty. PSA vehicles decoded
since the catalog module landed (72c0de6) showed categories but 0 parts
(41/43 affected). Exclude /psa/ from isP4LegacyPath so these paths reach
fetchPsaIllustrations / fetchPsaParts. Heals existing vehicles on demand;
no re-decode needed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 00:28:41 +03:00
078076b619 feat(demo): public /demo namespace serving pre-warmed VW Golf 2003 catalog
Replaces the old marketing "guided tour" /demo with a real, fully-functional
catalog browsing experience for the pre-warmed example vehicle. No auth
required, no upstream calls — entirely served from prod DB.

Backend (apps/api/src/demo):
* New @Public() controller exposing five endpoints under /api/demo:
  - GET /vehicle                    → demo vehicle metadata
  - GET /categories/tree            → top-level category tree
  - GET /categories/search?q=       → cross-tree search
  - GET /categories/:id             → getCategoryWithParts (parts+schema+hotspots)
  - GET /categories/:id/children    → drill children
* DemoService validates every category id against DEMO_VEHICLE_ID before any
  downstream service call — the public surface can't be used to read an
  arbitrary vehicle's catalog (1-row SELECT, NotFound on miss or wrong owner).
* Vehicle id is env-driven (DEMO_VEHICLE_ID, defaults to the pre-warmed
  WVWZZZ1JZ3W597935 — VW Golf 2003 with 277 cats / 9841 parts / 178 schemas
  fully drilled in prod).
* Wires CategoriesModule (already exports CategoriesService) — zero new
  business logic, just a thin public façade.

Frontend (apps/web):
* /demo (replaces old marketing page): vehicle header + top categories grid
  reading /api/demo/* + sticky DemoBanner with sign-up CTA.
* /demo/categories/$categoryId: drill page rendering either a children grid
  (parent) or the existing SchemaViewer + parts panel (leaf) — same shape
  the dashboard uses, so hotspot overlay, breadcrumb trail, retry on
  upstream loadError all just work.
* DemoBanner: sticky top, "Örnek araç: {label} — Kayıt Ol" CTA. The
  "Yeni VIN sorgula" explicit paywall trigger lands in a follow-up task.
* PostHog events: demo_loaded (source query-param-aware),
  demo_category_clicked, demo_category_detail_viewed, demo_to_register_click
  (banner / footer / category_footer placements).
* usePageMeta gains an opt-in `noindex` flag — demo sets it to noindex,follow
  for the first 4-6 weeks per spec; cleaned up on unmount so SPA navigation
  doesn't carry it to the next route.
2026-06-02 00:07:29 +03:00
286307155e fix(pcat-auth): two-armed post-goto poll — fail-fast on dead sites
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
The post-goto token poll ran a blind 20s wait regardless of whether
page.goto succeeded or threw. On a healthy goto the widget API call
fires within ~1-2s; on a failed goto the request either already went
through (rare) or never will (common). The 20s cap was the dominant
cost on failed-site attempts — verified tonight as a 28s "No token
after ..." log on auto-komplekt after page.goto ERR_TIMED_OUT.

* CAPTURE_POLL_AFTER_OK   = 10 (5s)  — token usually arrives in <2s
* CAPTURE_POLL_AFTER_FAIL = 4  (2s)  — brief grace then bail

Per-attempt worst case on a dead site: 10s goto + 2s grace = 12s
(was 10s + 20s = 30s). On a healthy site, well-known capture times
(3-5s) stay comfortably inside the 5s post-goto cap.
2026-06-01 22:17:41 +03:00
dcf7e068c9 feat(pcat-auth): cache JWT slot in Redis (PL24-style cross-restart hydration)
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
A captured slot now lives in Redis under `pcat:jwt:slot` for the JWT's
remaining lifetime (minus a 60s safety buffer). On module init we try
Redis before launching Playwright — if a fresh slot is there we adopt it
and schedule its refresh, skipping the ~5s capture entirely. After every
successful capture+validation we publish to Redis so the next restart (or
any sibling pod) can inherit. invalidateSession deletes the Redis copy
because a 401/403 means the cached IP-binding is dead.

Token is still IP-bound to its proxyPort. If a hydrating container reads
the slot but the proxy has rotated away from the captured IP, the next
upstream call 401s and the existing invalidateSession fallback re-captures
locally — so worst case = today's cold-capture behavior, never worse.

Note: dev and prod use separate Redis instances. This patch reaches PL24
parity (same-env redeploy hydration); a true dev↔prod shared cache would
need either an external Redis or an internal-token bridge.
2026-06-01 22:08:45 +03:00
30be00465b fix(pcat-auth): cap concurrent captures at JWT_SITES.length, not 3
Concurrent captures are bound by how many distinct partner sites we can
drive in parallel — each capture needs its own site so they don't collide
on the synchronous siteIndex round-robin. Ports are effectively unlimited
(10000-10999) and Playwright contexts are isolated, so the previous
arbitrary cap of 3 left headroom on the table when the pool is cold and
N>3 user clicks race in at once.

Tying the cap to JWT_SITES.length also means future site additions or
removals auto-adjust the ceiling.
2026-06-01 19:58:35 +03:00
9627e58cdd fix(pcat-auth): allow 3 concurrent captures + validate slots upstream
Two follow-ups to the warm-pool patch:

* Semaphore 1 → 3. Playwright contexts are isolated, the captureToPool
  site/port allocation is synchronous (no race), and concurrent user
  clicks that all miss the pool no longer serialize behind a single
  ~5s capture. Peak memory grows from one context to three; each is
  short-lived.

* Capture-time validation. After Playwright extracts the JWT, do one
  cheap upstream call (/car/info with the public demo VIN) through the
  same proxy port before pushing the slot to the pool. DataImpulse
  occasionally rotates to IPs the partner widget can load but the
  upstream API can't reach, or that get instantly 401/403'd; those
  ports used to spend 30s timing out on the first real user click.
  Failures rotate to the next site within the existing 4-retry budget.

Adds ~1s to each successful capture; saves up to 30s per dead slot.
2026-06-01 19:50:19 +03:00
a4644463e4 fix(pcat-auth): keep JWT pool warm 24/7 + tighten capture timeouts
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
The pool was gated to 09:00-19:00 Istanbul so off-hours the slot list was
empty and every first user click paid a 30-120s cold-start tax (Playwright
nav + retries across 4 sites). On the dashboard category page this looked
like an infinite skeleton loader. Pool refresh is cheap (~160 captures/day
per slot) — keep it warm always.

* isBusinessHours() now returns true; scheduleBusinessHours() is a no-op
  stub so the onModuleInit call site and businessHoursTimer field stay
  valid. All 8 existing gates (initial capture, replacement, refresh
  scheduling, refresh skip, scaling) become unconditional.
* getIstanbulTime() dropped — last consumer gone.
* PAGE_TIMEOUT 30s → 10s. Healthy partner sites load in <5s through the
  DataImpulse proxy; the longer wait only stretched dead-port retries.
* Drop e-acca.com from JWT_SITES — the current DataImpulse rotating proxy
  (74.81.81.81:10000-10999) cannot reach it; page.goto always blocked
  until PAGE_TIMEOUT instead of failing fast like the other sites.

Worst-case capture wall-clock: ~120s (4 × 30s) → ~40s (4 × 10s).
First off-hours request, with warm pool: ~120s+ → instant.
2026-06-01 19:26:44 +03:00
3613caa072 fix(catalog-source): gate emex parts behind allowlist (safety)
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
The catalog-wide bridge in EmexSourceDbService.fetchCategoryParts was
measured against vehicle_parts on 2026-06-01 and found to return 7-114x
more parts than belong to the requesting vehicle, with 49-98 wrong OEM
codes per 100 served. That directly violates the project rule that the
user must never see a wrong OEM.

Per-catalog noiseRatio sample (catalog-wide / per-vehicle):
  RENAULT201910 51x | FFIAT84 45x | VOLVO201410 24x | MB201810 14x
  AU1587 8x | BMW202501 70x (+ gid namespace mismatch ETK vs numeric)
  GM_C201809 114x | MINI202501 12x | LRE201412 7x | MAZDA2020 54x
  GM_OP201809 dump has only 1 wildcard vehicle (unique_key="_") so the
  single Crossland X "owns" all 47k Opel parts — same firehose served
  to any Opel sub-model in sase prod.

All alternative bridges were proven dead:
  SSD eşleştirme        - session-bound, 0/91 sase SSDs match dump
  scrape_queue_v2.vehicle_ssd - same session SSD format
  api_cache replay      - table empty (0 rows)
  wizard_parameters     - table empty (0 rows)
  VIN direct            - no VIN column in dump
The only viable per-vehicle bridge is vehicles.unique_key reconstruction
from raw_data.parsedOptions, but sase currently stores the required 4
wizard fields on just 5/103 emex vehicles (all Renault). That work is
follow-up; this patch only stops the bleeding.

Change:
- Add EMEX_SOURCE_DB_ALLOWED_CATALOGS env (comma-separated, default "")
- EmexSourceDbService.fetchCategoryParts returns null unless catalogCode
  is in the allowlist. Empty allowlist = service is effectively off for
  parts, full fallthrough to live emex.
- Connection pool stays alive so the follow-up per-vehicle bridge /
  schema-only path can use it without flipping env.
- Boot logs warn loudly when connected with an empty allowlist.

Prod was never affected — CATALOG_SOURCE_DB_ENABLED was unset there. This
fixes dev branch behaviour (default-on since commit 3a3a7d3) and keeps
prod safe by default once main is promoted.

Files:
- packages/config/src/index.ts        env schema + audit notes
- apps/api/src/config/configuration.ts parse allowlist into string[]
- apps/api/src/integrations/catalog-source-db/emex-source-db.service.ts
  allowlist field, init logging, fetchCategoryParts gate, class doc
- docker-compose.coolify.yml          env injection for api + worker

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-01 18:39:59 +03:00
Süper Panel
778931c880 fix(observability): whitelist Sentry ingest in CSP connect-src
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
The browser Sentry SDK initialised fine (DSN reached the bundle,
__SENTRY__ carrier registered) but envelope POSTs were silently
blocked by the existing CSP — `connect-src` didn't list any Sentry
host. Playwright verification on dev.sase.tr confirmed zero requests
to *.sentry.io even after a deliberate uncaught error.

Adds https://*.ingest.de.sentry.io (otolog org lives in the EU/de
region; this matches both the python and sase-web project DSNs).
2026-06-01 18:28:47 +03:00
3a3a7d3ebf chore(catalog-source): split per-source kill switches; default pcat off
Verified 2026-06-01 against dev's 103 unique pcat carIds: the current pcat
dump's deep-scrape (7.978 cars with real parts data via schema_parts or
part_groups+part_group_items) targets a US/JDM-market subset — Toyota 2112,
Nissan 1508, Audi 1311, Chevy 1050, Hyundai 745. **None** of sase's TR-market
vehicles intersect that rich subset:

  - 18/103 sase carIds are in dump.cars at all (registry only)
  - 0/103 yield parts via Bridge A (schema_images → schema_parts)
  - 0/103 yield parts via Bridge B (part_groups → part_group_items)

Even the cars that match by exact carId (Fiat Doblo 368 schemas, Renault
Megane, Bravo 456 schemas) have only diagram metadata — no parts annotation.
The dump scraper finished tier-1 (catalog/model/car listing) and tier-2
(schema diagrams) for these, but stopped before tier-3 (parts annotation).

Under the strict "always correct OEM" constraint there is no safe pcat lookup
today. Disable it. The container stays up for future use cases (OEM cross-
reference search, alt-part matching) and so we can flip the env back without
a code change if a richer dump arrives.

EMEX stays on (its catalog-allowlist is the next step).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 17:50:07 +03:00
94117dd75d chore(api): bump source-db hit logs to log level for dev verification
Easier to see [source-db hit pcat/emex] lines while dev verifies coverage.
Can be downgraded back to debug once we've measured prod hit rates.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 15:07:47 +03:00
41acff2ab2 fix(api): emex source-db lookup uses catalogCode+gid, not SSD
Initial design routed emex lookups through vehicles.rawData.ssd → dump vehicles
→ vehicle_parts. Smoke test against prod ssd values: 0 / 10 matched. EMEX
regenerates the SSD on every decode session, so sase's stored SSD never
matches the SSD the dump scraper recorded for the same physical vehicle.

Pivot to a catalog-wide bridge that actually works:
  catalogs.code    ↔ vehicles.rawData.catalogCode  (e.g. "RENAULT201910")
  part_groups.group_id ↔ categories.externalId    (e.g. "11754")
  → parts via vehicle_parts.group_id (dump's parts.group_id is 100% NULL)

Verified coverage on prod's 8287 unique (catalogCode, gid) pairs: 25/26
catalog codes resolve, 7178 pairs hit a part_group (87%), 5919 of those
return actual parts via vehicle_parts (~71% net). Tradeoff: returns all
parts in the (catalog, group) across every variant in the catalog, so the
result is slightly noisier than the live per-vehicle scrape. Acceptable —
parts overlap heavily and the upstream-call savings outweigh the noise.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 15:06:48 +03:00
4724a71113 feat(api): source-DB lookup-first for pcat/emex catalog fetches
Adds an optional local-dump lookup layer in front of the live PartsCatalogs and
EMEX scrapes. When enabled, getCategoryWithPartsInner queries a Postgres
(pcat) or MariaDB (emex) dump for the requested schema/group's parts and
hotspots; on miss it falls through to the existing upstream call unchanged.
Hits avoid the live API, its cooldown, and its rate-limits — direct DB latency.

- New CatalogSourceDbModule with PcatSourceDbService + EmexSourceDbService
  (raw SQL, no Drizzle schema modeling — dump shapes are frozen snapshots).
- pcat lookup keys on schema_images.schema_ext_id (the dump's column that
  matches sase's pcat groupId; observed ~7% hit rate on prod's 6596 unique
  groupIds, of which ~10% have schema_parts → ~3-5% net parts coverage).
  Joins schema_parts → parts directly; the dump's part_groups+part_group_items
  linkage covers 0 of our hits, so we skip that path entirely.
- emex lookup uses (catalog_id, ssd) → vehicles.id then (vehicle_id, group_id)
  → vehicle_parts → parts + part_images. The ssd is already persisted into
  vehicles.rawData.ssd by the existing emex.mapper, no extra capture needed.

Gated behind CATALOG_SOURCE_DB_ENABLED + PCAT_SOURCE_DB_URL / EMEX_SOURCE_DB_URL.
All three default unset, so this commit is a no-op until prod env is configured.
Adds mysql2 dep for the MariaDB client.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 00:35:45 +03:00
1d9f59ba86 fix(api): pause whole worker on cooldown, only per-job defer for off-hours
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Per-job moveToDelayed for cooldown livelocked the worker: the activity key is
refreshed by every user request, so 60s later the deferred job comes back, key
is still set, defers again. Last 3h on prod logged ~1800 deferrals against 12
real inits and one completion every ~6 min.

Tag RateLimitError with a `cause`. checkCooldown throws "cooldown"; the worker
now calls `this.worker.rateLimit(delayMs)` and throws Worker.RateLimitError() —
the whole queue waits once instead of cycling every job. checkTimeWindow throws
"time-window"; that branch keeps the existing job.moveToDelayed (per-job) so
EMEX (no scrape window) keeps flowing while PL24/pcat jobs sleep till 09:00.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 12:17:05 +03:00
f733186a04 fix(csp): allow destek.sase.tr for the Chatwoot widget
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
The live-chat SDK, widget iframe, websocket and avatars are served from
destek.sase.tr; the helmet CSP didn't whitelist it, so the browser
blocked sdk.js (script-src violation) and the widget never loaded.
Add destek.sase.tr to script-/img-/media-/connect-/frame-src (+wss for
ActionCable).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 19:50:18 +03:00
1787607be5 feat: embed Chatwoot live-chat widget (destek.sase.tr)
Site-wide live-chat widget served from the self-hosted Chatwoot at
destek.sase.tr, with verified user identity and vehicle context.

- apps/web: lib/chatwoot.ts loads the SDK lazily (mirrors the PostHog
  init pattern), init in main.tsx, identify logged-in users in __root
  via a server-computed HMAC, and attach the viewed vehicle (VIN/brand/
  model) as contact custom attributes on the vehicle detail page.
- apps/api: GET /api/chatwoot/identity (AuthGuard-protected) returns
  HMAC-SHA256(user.id) so the widget can use verified identity.
- env: VITE_CHATWOOT_BASE_URL + VITE_CHATWOOT_WEBSITE_TOKEN (build-time,
  wired through docker-compose.coolify.yml build args + Dockerfile ARG)
  and CHATWOOT_HMAC_TOKEN (api runtime). All optional — widget and
  endpoint no-op when unset.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 19:28:50 +03:00
c7e59b90e0 fix(api): defer rate-limited prefetch jobs instead of failing them
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
The worker threw a custom RateLimitError for active-user cooldown and off-hours
windows, but BullMQ treated that as a normal failure: it retried with the
queue's exponential backoff (ignoring retryAfterMs) and dropped the job to
"failed" after attempts:3. So during busy hours sub-jobs were discarded en
masse and the requested defer time (e.g. "wait until 09:00") never applied.

Convert RateLimitError into job.moveToDelayed(now + retryAfterMs, token) +
DelayedError — BullMQ's contract for "not done, not failed, retry later". This
does not consume an attempt and honours the exact delay, so cooldown jobs wait
~60-120s and off-hours jobs wait until the window reopens, then resume. Genuine
errors still fail/retry normally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 18:02:36 +03:00
dae76085ef fix(api): translation job IDs must not contain ':'
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Same BullMQ 5.68 restriction as the prefetch fix (acd5691): a custom job ID
containing ':' is rejected with "Custom Id cannot contain :". enqueueTranslation
built jobId `tr:<base64>`, so every enqueue threw — and both call sites
fire-and-forget with .catch(warn), so it failed silently: fresh terms were
NX-flagged as queued but never actually enqueued, leaving new EMEX/PCAT terms
untranslated (English). Use 'tr-' prefix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 16:49:46 +03:00
acd5691f6b fix(api): prefetch sub-job IDs must not contain ':'
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
BullMQ 5.68 rejects custom job IDs containing ':' (its key separator) with
"Custom Id cannot contain :". prefetch-init's addJob built sub-job IDs as
prefetch:<vehicleId>:<categoryId>:<action>, so every attempt to queue a
children/parts job threw and the whole init failed. This was latent in the
reactive path (failures just logged) and surfaced once the hourly backfill
started driving inits at volume. Use '-' as the separator; the IDs only need
to be deterministic for dedup, not parseable.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 15:18:04 +03:00
5f0bdb467d fix(api): gate catalog backfill on prod host, not NODE_ENV
dev.sase.tr (staging) and sase.tr (prod) BOTH run NODE_ENV=production with
SEPARATE databases, so the previous NODE_ENV check would have let the hourly
backfill sweep run against the dev DB too. Gate on the canonical prod host
instead (COOLIFY_FQDN / BETTER_AUTH_URL), with an explicit
CATALOG_BACKFILL_ENABLED override. Default off for any unknown host.

New isCatalogBackfillEnabled() helper used by both the cron registration and
processBackfillScan; dev redeploy now removes the stale scheduler.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 14:29:07 +03:00
d0ee4df57f feat(api): hourly catalog backfill for decoded-but-unfetched vehicles
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
prefetch was reactive-only (on decode) and processInit read top-level
categories straight from DB, so a vehicle decoded but never viewed got
no catalog. Add a production-only sweep so no decoded vehicle is left
without catalog data.

- processInit self-seeds top categories via getCategoryTree when DB has
  none (fetches+inserts top groups from PL24/PSA/EMEX), closing the
  never-viewed gap for both reactive and backfill paths
- new backfill-scan job + hourly cron: phase 1 queues decoded vehicles
  with zero parts, phase 2 rolling createdAt-cursor rescan of all decoded
  vehicles (prefetch-init is idempotent → gap-fills partial ones)
- guardrails: skip wave if queue backlog > 1000, per-source cooldown,
  business-hours window (isWithinTimeWindow), batch <=20, in-flight guard
- PRODUCTION ONLY: gated on NODE_ENV both at cron registration and in
  processBackfillScan; dev has a separate DB and must not scrape

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 14:04:32 +03:00
8718b08415 feat(catalog): full-catalog search on vehicle page (leaf categories + OEM parts)
The vehicle page search previously only filtered category names at the
currently rendered level. Add a server-side cross-tree search over what's
already drilled into the DB.

New GET /categories/search/:vehicleId?q= returns two sections:
- categories: name-matched leaves UNION the leaf categories that contain a
  matching part (with hit count). "fren balatası" matches no leaf by name —
  the pads are parts under leaves like "Disk freni" — so the union surfaces
  the right leaves.
- parts: parts matching every token on name, or the raw query on oem_code,
  with OEM + leaf + breadcrumb.

Pure DB read (no upstream drill); a treeIncomplete hint is returned when the
vehicle's tree looks barely drilled. Frontend adds a debounced search box on
the vehicle page that hides the normal browse while active.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 13:30:21 +03:00
2cc2a917a5 fix(api): CSP'ye Cloudflare Turnstile'ı ekle (widget engelleniyordu)
helmet CSP script-src/connect-src/frame-src challenges.cloudflare.com'u
engellediğinden Turnstile widget'ı yüklenemiyordu → token üretilmiyor →
captcha enforce edilen login/register/contact "şifre hatalı" ile reddediyordu.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 02:21:15 +03:00
7eb2e85bbc feat(auth): "beni 30 gün hatırla" + login Turnstile koruması
- session.expiresIn = 30 gün; login'de rememberMe checkbox signIn.email'e bağlandı
  (işaretli=30 gün kalıcı çerez, değilse oturum çerezi)
- better-auth captcha plugin'i (cloudflare-turnstile, TURNSTILE_SECRET_KEY varsa)
  sign-in/sign-up'ı korur; login formuna Turnstile widget'ı + x-captcha-response

Not: auth.ts ve login.tsx her iki özelliği birlikte içerdiğinden tek commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 02:14:18 +03:00
a09182fb98 feat: Cloudflare Turnstile (register + contact) + captcha altyapısı
- Turnstile widget bileşeni (public site key gömülü, VITE_TURNSTILE_SITE_KEY ile override)
- register: signUp.email'e x-captcha-response header'ı
- contact: token body'de; ContactService Cloudflare siteverify ile doğrular
  (TURNSTILE_SECRET_KEY yoksa atlanır), contact.dto'ya turnstileToken
- @sase/config: TURNSTILE_SECRET_KEY env (opsiyonel)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 02:14:06 +03:00
1c79c18aff feat(api): wire Novu lifecycle email triggers
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
Route all lifecycle/transactional emails through Novu
(api.bildirim.semih.ai, delivered via Postal). A framework-agnostic
client is shared by the NestJS API and the standalone BullMQ worker.

- welcome + referral on signup (better-auth user.create.after)
- email-verification + password-reset (auth.ts; token links never
  track-wrapped so the one-time token survives)
- referral-qualified / referral-reward to the referrer on qualification
- payment-success / payment-failed in the Stripe webhook handlers
- trial-ending + win-back via a new daily lifecycle-email cron (worker),
  idempotent via a 1-day endDate window (no sent-flag column)
- signed track.sase.tr CTA links when MAILTRACK_SECRET is set
- NOVU_* / APP_PUBLIC_URL / MAILTRACK_SECRET env added to config,
  validation, .env.example and both compose service blocks

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 01:30:39 +03:00
c633e846e2 fix(payments): curate /payments/me response — stop leaking internal columns, join plan name
getMyPayments returned the raw payment row, exposing internal fields
(adminNote, iyzicoPaymentId, bankAccountId, session/intent ids) to the
end user. Replace with an explicit projection that returns only what the
billing UI needs, joins planName from the subscription's plan (was always
"-"), and surfaces Stripe receipt availability as a hasStripeReceipt
boolean instead of the raw payment intent id. Frontend reads the boolean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 00:31:33 +03:00
1b9492ba49 style(api): drop non-null assertion in refund amount
Replace `input.amount!` with `input.amount ?? Number(payment.amount)`,
which is behaviour-identical (undefined amount = full refund = full
amount) but satisfies lint/style/noNonNullAssertion.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 00:15:41 +03:00
7a8fe5e98e feat(payments): Stripe receipt link on billing rows
Add GET /payments/:id/receipt — resolves the Stripe-hosted receipt URL
from the payment intent's latest charge (ownership-scoped; returns EFT
receipt directly when present, null otherwise). Wire a "View receipt"
action on completed Stripe rows that fetches the URL on demand and opens
it, with a toast when none is available.

Closes the last billing-audit item (#8).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 00:14:28 +03:00
57b3cb7ebe feat: contact sayfasına iletişim formu + POST /contact endpoint'i
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
- Yeni contact modülü: @Public POST /contact, zod validation, @Throttle 5/10dk
  spam koruması; EmailService ile admin@sase.tr'ye mail (reply-to = gönderen),
  kullanıcı girdileri HTML-escape
- EmailService: replyTo desteği eklendi
- contact.tsx: useState + zod + @sase/ui ile iletişim formu (mevcut form
  pattern'iyle tutarlı; yeni form kütüphanesi yok), toast + alan validasyonu

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 01:27:56 +03:00
77de860211 feat: VIN textbox WMI marka ikonu + WMI haritasını shared'e taşı/genişlet
- WMI_BRAND_MAP + getBrandFromWmi packages/shared'e taşındı (tek kaynak);
  corgi.service artık buradan import ediyor (davranış aynı, testler geçiyor)
- VinBrandIcon: VIN'in WMI'ı bilinen markaya denk gelince büyüteç yerine
  marka logosu pop animasyonuyla görünür (landing + search VIN textbox)
- WMI haritasına prod DB'de decode edilmiş VIN'lerden 23 eksik WMI eklendi
  (Türkiye fabrikaları NM4/NMT/NLA/NLH/NMB dahil; tüm DB WMI'ları artık tanınıyor)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 23:51:40 +03:00