fix(insights): tighten session fingerprint so duplicate UX issues dedupe #5
Reference in New Issue
Block a user
Delete Branch "fix/insight-fingerprint-dedupe"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
ℹ️ Stack notu
Bu PR
fix/insight-vin-validation-distinction(PR #4) branch'i üzerine kuruldu çünkü PR #4 henüz merge edilmemişti. Sonuç: PR #5 hem PR #4'ün düzeltmesini hem de fingerprint dedupe düzeltmesini içeriyor. İki ayrı commit olarak görünür:9a479f9— fix(insights): distinguish client-side VIN validation rejects from provider failures (PR #4 ile aynı)a7fe80f— fix(insights): tighten session fingerprint so duplicate UX issues dedupe (bu PR'ın asıl içeriği)Yani PR #5 merge edilirse PR #4'e gerek kalmaz — PR #4 manuel olarak kapatılabilir.
Sorun
Aynı kök sebepli UX problemleri panel'de tek bir Insight olarak dedupe edilmiyor — son 1 günde 6 farklı
newinsight açılmış, hepsi "kategori ağacında bouncing → parts panel'i bulamama → rage click" senaryosunu farklı kelimelerle anlatıyor:cmpclg7b9000dflajkg27hxvtcmp5yj1pa00071bakh9nqdh22cmpwrb82s002f14fza8lbc7f4cmpwqgbyc002d14fzhnumfatwcmpwlwt4a002314fzg8wyi3nhcmpoeec4l001j14ozjjov5vrmSonuç: tetikçi (issue açılması, Telegram alert) sadece ilk bir-iki insight için ateşleniyor, kalanlar arka planda birikiyor; aynı problem için "1 insight" diyince aslında 6'sı saklı kalıyor.
Kök sebep
apps/worker/src/lib/compress.ts:477— session fingerprint inputs:İki sorun:
vehicleIdvecategoryIdher user için farklı → her session benzersiz fingerprint üretiyor →prisma.insight.findUnique({projectKey_fingerprint})cache miss → yeni insight açılıyor.header.severityfingerprint'in parçası — aynı kök sebep, farklı rage-click yoğunluğu yüzünden bazen P1, bazen P2/P3 çıkıyor. Severity bucket'ın bir özelliği olmalı, kimliği değil.Düzeltme
apps/worker/src/lib/compress.ts:normalizePath()helper: UUID v1–v8, ULID, CUID2, sayısal id path segmentlerini:idile değiştirir. Anlamlı path kelimelerine dokunmaz.header.severityçıkarıldı.failedEndpointsda normalize ediliyor (örn./api/vehicles/123/parts→/api/vehicles/:id/parts).Tag set + normalized path + first-error + first-failed-endpoint yeterli ayrışım sağlıyor çünkü farklı UX problemleri zaten tagger'da farklı tag'ler alıyor (
vin_decode_*,payment_*,search_validation_*, …).Validation
apps/workertest runner kullanmıyor; smoke/tmp/check_fingerprint.tsile yapıldı (12/12 ✓):pnpm -F worker typechecktemiz.Etkisi
insightssatırları olduğu gibi kalır — eski yakın-duplicate insight'lar panel UI'dan elle merge edilmeli (founderNotes + status=dismissed).analyze.ts:79'dakiprisma.insight.findUnique({projectKey_fingerprint})cache lookup mantığı aynen çalışır.Deploy
Merge → main → panel-worker otomatik redeploy.
Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com
🤖 Generated with Claude Code
Reproducing insight cmpvfrjgc000114fzc7cdyh66 (a P1 "PL24 timeout" false positive): trial user typed VW part numbers ("500 907 521", "5Q0 907 521") into the VIN field on /; client-side regex rejected them with "Geçersiz şase numarası. 17 karakter olmalı". No provider was called. The pipeline still tagged the session as `vin_decode_fail_pattern`, routed to `provider_quality`, and the LLM dutifully invented a PL24 outage. Root cause spans three files: 1. tagger.ts grouped vin_decode_failed by `provider_attempted ?? source`. When `provider_attempted` is missing, `source: "landing"` (a UI location) was treated as a provider name, so a 1-provider set was synthesized and `vin_decode_fail_pattern` (P1) was emitted. 2. compress.ts formatCustom whitelist excluded `error`, `source`, `vin`. The LLM therefore never saw "Geçersiz şase numarası" or the offending input. Pattern 3 mechanical hypothesis told it "check provider health" regardless. 3. prompts.ts pickPromptTag routed any `vin_decode_fail_pattern` straight to `provider_quality` with no input-quality check, and the v3 system prompt had no guardrail for client-side validation rejects. Fix: - tagger: detect client-side rejects by `error` regex (Turkish + English) and by VIN shape (length != 17 or contains I/O/Q). When all fails are client rejects, emit new tag `vin_decode_client_validation_fail` at P3 instead of `vin_decode_fail_pattern` at P1. Real provider failures now require `provider_attempted` to be set (no more `source` fallback). - compress: add `error`, `source`, `vin` to the formatCustom property whitelist so the LLM can see the actual failure context. Split Pattern 3 into client-reject vs. real-provider-failure branches with distinct Turkish hypotheses. - prompts: route `vin_decode_client_validation_fail` to `ux_friction` before the provider rule. Ship provider_quality v4 with an explicit guardrail instructing the model to return confidence ≤0.15 and reclassify when the inlined event properties show client-side rejection. The seed-runtime upsert path deactivates the active v3 template on next worker boot and inserts v4 in its place — no manual SQL needed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>Same root cause was producing one Insight row per user because URL carried vehicleId/categoryId UUIDs and the raw URL fed the fingerprint hash. Result: 6 active "new" insights describing the same kategori-bouncing problem with slightly different LLM phrasing, none deduplicated, only one (cmpclg7b9 → sase.tr#76) had been triaged. apps/worker/src/lib/compress.ts: - new normalizePath() — collapses UUID / ULID / CUID2 / numeric path segments to ":id", conservative on plain words. Mirrors what the LLM already sees in the timeline. - fingerprint inputs now: tags (sorted) | normalizePath(url) | normalizeError(errors[0]) | normalizePath(failedEndpoints[0]) - removed header.severity from the hash — severity is a property of the Insight bucket, not its identity; rage-click counts pushing the same root cause across P1/P2/P3 was forcing extra rows. Smoke (12/12 pass via tsx /tmp/check_fingerprint.ts): - 6 historical kategori-bouncing URLs → 1 fingerprint - severity changes don't move the hash - distinct tag sets / distinct errors still split Historical rows are untouched — only new compressed sessions get the new hash. Old near-duplicate insights can be merged manually via the panel. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>