#95 — Vorratsschrank — KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2) #95

Closed
opened 2026-08-18 13:13:56 +02:00 by lena · 3 comments
lena commented 2026-08-18 13:13:56 +02:00 (Migrated from git.butzei.de)

Story: Vorratsschrank — KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2)

As a Haushaltsmitglied,
I want to ein Produkt, dessen Barcode in keiner Datenbank bekannt ist, per Foto oder Spracheingabe statt
reiner Texteingabe identifizieren lassen,
so that ich unbekannte Produkte nicht jedes Mal komplett von Hand eintippen muss.

Depends on: #94 (Phase 1 — Barcode-Scan, manuelle Text-Eingabe als Fallback existiert bereits).

Acceptance criteria:

  • Ist ein gescannter Barcode weder in der öffentlichen Produktdatenbank noch in der eigenen
    Wissensdatenbank bekannt, bietet die App zusätzlich zur bestehenden Text-Eingabe zwei weitere Eingabewege
    an: Foto und Spracheingabe.
  • Foto: ein aufgenommenes/hochgeladenes Bild wird an eine externe KI geschickt, die einen
    Produktnamen-Vorschlag liefert; der Nutzer bestätigt oder korrigiert ihn.
  • Spracheingabe: eine gesprochene Beschreibung wird transkribiert und als Produktname-Vorschlag übernommen,
    ebenfalls bestätigt/korrigierbar.
  • Das bestätigte Ergebnis wird wie in Phase 1 dauerhaft mit dem Barcode verknüpft gespeichert (nie
    zweimal gefragt).

Out of scope for this story:

  • Auswahl des konkreten KI-Anbieters/Modells — technische Entscheidung des Architects, ggf. in Abstimmung mit
    dem Menschen (Kosten, Datenschutz).
  • Offline-Fähigkeit: anders als #82 (rein lokal im Browser) ist eine externe Anbindung hier bewusst akzeptiert,
    da lokale Modelle für generische Foto-/Spracherkennung in dieser Qualität im Browser aktuell nicht realistisch
    sind.

Sicherheits-Vorprüfung ist für diese Story verpflichtend, nicht optional: Dies wäre die erste Funktion der
App, die Nutzerdaten (Fotos aus der eigenen Küche, Sprachaufnahmen) an einen externen Drittanbieter sendet.
Bevor implementiert wird, muss die Security-Agent-Vorprüfung (siehe ai/roles/00_team_overview.md,
Feature-Zyklus) explizit klären: welcher Anbieter, welche Daten die App verlassen, ob sie beim Anbieter
gespeichert/für Training verwendet werden, und ob vor der ersten Nutzung ein Zustimmungs-/Datenschutzhinweis
nötig ist.

Open questions:

  • Gibt es eine bevorzugte KI-Anbindung (z. B. bereits vorhandener Account/API-Key), oder soll der Architect
    frei vorschlagen? → Rückfrage an den Menschen bei Ausarbeitung.
    Resolved by human, 2026-08-11:
    "Please use an common standard, and make the url and api key configureable. It should work with local
    ollama as well." → OpenAI-compatible HTTP API (/chat/completions for vision, /audio/transcriptions
    for speech), base URL + API key + model names configurable via deployment config (env vars/
    appsettings.json, same pattern as SMTP/push VAPID), not a per-user setting. See
    95_pantry_ai_product_recognition_phase2_design.md for what this does and doesn't cover (Ollama serves
    the vision half natively; its own text-generation server has no audio-transcription endpoint, so speech
    needs a different self-hosted or cloud endpoint that speaks the same Whisper-compatible API shape).
# Story: Vorratsschrank — KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2) **As a** Haushaltsmitglied, **I want to** ein Produkt, dessen Barcode in keiner Datenbank bekannt ist, per Foto oder Spracheingabe statt reiner Texteingabe identifizieren lassen, **so that** ich unbekannte Produkte nicht jedes Mal komplett von Hand eintippen muss. **Depends on:** `#94` (Phase 1 — Barcode-Scan, manuelle Text-Eingabe als Fallback existiert bereits). **Acceptance criteria:** - [ ] Ist ein gescannter Barcode weder in der öffentlichen Produktdatenbank noch in der eigenen Wissensdatenbank bekannt, bietet die App zusätzlich zur bestehenden Text-Eingabe zwei weitere Eingabewege an: **Foto** und **Spracheingabe**. - [ ] Foto: ein aufgenommenes/hochgeladenes Bild wird an eine externe KI geschickt, die einen Produktnamen-Vorschlag liefert; der Nutzer bestätigt oder korrigiert ihn. - [ ] Spracheingabe: eine gesprochene Beschreibung wird transkribiert und als Produktname-Vorschlag übernommen, ebenfalls bestätigt/korrigierbar. - [ ] Das bestätigte Ergebnis wird wie in Phase 1 dauerhaft mit dem Barcode verknüpft gespeichert (nie zweimal gefragt). **Out of scope for this story:** - Auswahl des konkreten KI-Anbieters/Modells — technische Entscheidung des Architects, ggf. in Abstimmung mit dem Menschen (Kosten, Datenschutz). - Offline-Fähigkeit: anders als `#82` (rein lokal im Browser) ist eine externe Anbindung hier bewusst akzeptiert, da lokale Modelle für generische Foto-/Spracherkennung in dieser Qualität im Browser aktuell nicht realistisch sind. **Sicherheits-Vorprüfung ist für diese Story verpflichtend, nicht optional:** Dies wäre die erste Funktion der App, die Nutzerdaten (Fotos aus der eigenen Küche, Sprachaufnahmen) an einen externen Drittanbieter sendet. Bevor implementiert wird, muss die Security-Agent-Vorprüfung (siehe `ai/roles/00_team_overview.md`, Feature-Zyklus) explizit klären: welcher Anbieter, welche Daten die App verlassen, ob sie beim Anbieter gespeichert/für Training verwendet werden, und ob vor der ersten Nutzung ein Zustimmungs-/Datenschutzhinweis nötig ist. **Open questions:** - ~~Gibt es eine bevorzugte KI-Anbindung (z. B. bereits vorhandener Account/API-Key), oder soll der Architect frei vorschlagen? → Rückfrage an den Menschen bei Ausarbeitung.~~ **Resolved by human, 2026-08-11:** "Please use an common standard, and make the url and api key configureable. It should work with local ollama as well." → OpenAI-compatible HTTP API (`/chat/completions` for vision, `/audio/transcriptions` for speech), base URL + API key + model names configurable via deployment config (env vars/ `appsettings.json`, same pattern as SMTP/push VAPID), not a per-user setting. See `95_pantry_ai_product_recognition_phase2_design.md` for what this does and doesn't cover (Ollama serves the vision half natively; its own text-generation server has no audio-transcription endpoint, so speech needs a different self-hosted or cloud endpoint that speaks the same Whisper-compatible API shape).
lena commented 2026-08-18 13:13:57 +02:00 (Migrated from git.butzei.de)

design (95_pantry_ai_product_recognition_phase2_design.md)

Design: #95 — Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2)

Architect note (2026-08-11). Human resolved the story's own open question: use a common,
self-hostable HTTP standard rather than a single hardcoded vendor SDK, with the endpoint, key, and model
names all deployment-configurable, and Ollama specifically supported. That points at the OpenAI-compatible
API shape
— the de facto standard every major self-hosted inference server (Ollama, LocalAI,
llama.cpp's server, vLLM) and every cloud vendor with an "OpenAI-compatible" mode already implements:
POST {base}/chat/completions for vision (image content parts), POST {base}/audio/transcriptions for
speech-to-text (Whisper-shaped multipart upload). One HttpClient, one base URL, one API key, two
independently-configurable model names.

Scope carried over from Phase 1

#94's ScanPantryProductBarcodeCommand already has the exact mechanism the story's 4th AC asks for
("Das bestätigte Ergebnis wird wie in Phase 1 dauerhaft mit dem Barcode verknüpft gespeichert, nie
zweimal gefragt"): when a barcode is unknown to both the pantry's own products (dedup by
(PantryId, Barcode)) and Open Food Facts, the caller supplies FallbackName and a new
PantryProductEntity permanently linked to that barcode is created. Phase 2 does not touch that
persistence path at all
— it only adds two new ways to produce the string the user was already typing
by hand: a photo-recognition suggestion and a speech-transcription suggestion, both still shown to the
user for confirmation/correction before the existing submit path runs. Same for Shopping's
ScanShoppingProductBarcodeCommand (#115 already shares Open Food Facts lookup with Pantry; this extends
that sharing to the new AI lookups too).

New deployment config — AiSettings

Ai:BaseUrl              e.g. https://api.openai.com/v1, http://ollama-host:11434/v1, or a self-hosted
                         Whisper-compatible server's URL. Empty = feature fully disabled.
Ai:ApiKey                Sent as `Authorization: Bearer {key}` when non-empty. Ollama ignores it; most
                         cloud vendors require it.
Ai:VisionModel           e.g. "gpt-4o-mini", "llava", "moondream", "minicpm-v". Empty = photo recognition
                         disabled even if BaseUrl is set.
Ai:TranscriptionModel    e.g. "whisper-1". Empty = speech recognition disabled even if BaseUrl is set.

Two independent bool flags, not one: IsPhotoRecognitionConfigured / IsSpeechRecognitionConfigured,
each requiring BaseUrl and its own model name. This is a deliberate design response to the human's
own "should work with local Ollama" requirement — Ollama's OpenAI-compatible layer serves
/chat/completions (including vision models) but has no /audio/transcriptions endpoint at all. An
operator running Ollama-only sets VisionModel and leaves TranscriptionModel blank; they get the photo
option only, not a "Sprache" button that always 501s. Same env/appsettings/docker-compose wiring pattern
as Email/Push (.env.example documents AI_BASE_URL/AI_API_KEY/AI_VISION_MODEL/
AI_TRANSCRIPTION_MODEL, mapped through docker-compose.yml's Ai__BaseUrl etc.). No
ValidateOnStart() — unlike AppSettings.FrontendBaseUrl, "all blank" is a valid, fully-supported
"feature off" state, not a misconfiguration.

New backend surface

  • IAiProductRecognitionClient (CqsTodo/Ai/), one HttpClient-backed implementation:
    • TryRecognizeProductFromPhoto(byte[] jpegBytes, CancellationToken) — one-shot chat completion, a
      fixed system prompt constraining the model to respond with a short grocery-product name and nothing
      else (no free-form chat, no injected user text — the only user-controlled input is the image itself),
      image sent as a data:image/jpeg;base64,... content part per the OpenAI vision message shape.
    • TryTranscribeSpeech(byte[] audioBytes, string contentType, CancellationToken) — multipart upload to
      /audio/transcriptions (file, model), returns the transcript text verbatim (the frontend dialog
      still shows it in an editable field — a mis-transcription is just a wrong prefill, not a data-integrity
      issue, same trust level Open Food Facts' suggestion already has).
    • Both: 30s timeout (vision/ASR inference is slower than Open Food Facts' 5s keyless lookup), both never
      throw out to the caller — any HTTP/timeout/malformed-response failure is logged and swallowed to
      null, matching OpenFoodFactsClient's own "external outage never breaks the feature" convention.
      Whichever modality's suggestion comes back null, the user still has the plain manual-entry text field
      as the ultimate fallback — nothing about this feature can make entering a product name harder than
      Phase 1 already made it.
  • Three new public requests (own file each, standalone — reusable by both Pantry and Shopping, not nested
    under either feature folder):
    • GetAiRecognitionCapabilitiesQuery() -> AiRecognitionCapabilitiesDto(bool PhotoSupported, bool SpeechSupported) — lets the frontend hide a button that would always fail instead of discovering that
      by calling it and getting null back.
    • RecognizeProductNameFromPhotoQuery(string ImageBase64) -> string?
    • RecognizeProductNameFromSpeechQuery(string AudioBase64, string ContentType) -> string?
    • All three WithAuthorization(_ => new AuthorizeIsCurrentUserAuthenticatedQuery()) — login-gated like
      push subscriptions/API keys, but deliberately not pantry/shopping-list-scoped, since recognizing
      "what's in this photo" needs no access to a specific list's data; the existing
      AuthorizePantryAccessForCurrentUserQuery/AuthorizeShoppingListAccessForCurrentUserQuery decorators
      still gate the actual Scan*Command that persists the confirmed name against a barcode.
  • Upload validation, shared with the existing avatar-upload path rather than re-implemented: extracted
    UpdateUserAvatarCommandHandler.ProcessUpload's decode/validate core (base64 parse, size cap, format
    allowlist via Image.Identify, dimension cap, animated-frame rejection, UnknownImageFormatException/
    InvalidImageContentException handling) into a shared ImageUploadValidator.LoadAndValidate(base64, maxBytes, maxDimension) returning a validated, decoded Image; each caller does its own
    encode/resize afterwards (avatar: crop-to-square 256px WebP; AI photo: cap-to-1024px-longest-side JPEG,
    no cropping — cropping a kitchen product photo could cut off the exact label text the model needs to
    read). Same two guards apply to the AI path as already apply to avatars: a decompression-bomb-shaped
    file is rejected before the full decode by dimension-checking by Image.Identify first, and the
    re-encode step means whatever EXIF/metadata (including GPS, if a phone attaches it) the original photo
    carried is never forwarded to the external AI endpoint — only re-encoded pixel data is.
  • Audio gets an equivalent but new, smaller validator (AudioUploadValidator) — no existing
    decode/re-encode path to extend, and no audio-processing library in this repo's dependency set, so this
    intentionally does not re-encode audio (there's nothing here to strip — MediaRecorder-produced
    WebM/Ogg has no EXIF-equivalent metadata risk). Enforces a size cap (10 MB — a few minutes of
    browser-recorded Opus easily fits, avoids a large-request DoS surface, similar reasoning to the photo's
    2 MB cap) and a magic-byte content-type sniff against an allowlist (WebM/Ogg/WAV/MP3 signatures) so a
    client-forged Content-Type claim can't smuggle an arbitrary file past the size check into the outbound
    multipart request to the external endpoint.

New frontend surface

  • New usePhotoCapture hook (ReactUi/src/hooks/), a sibling to useBarcodeScanner rather than an
    extension of it — found while implementing that by the time a user reaches the "unknown product"
    fallback, useBarcodeScanner's own successful-decode callback has already called controls.stop(),
    so there's no still-open stream left to snapshot from (the design's original plan to extend that hook
    didn't hold up against its actual lifecycle). usePhotoCapture opens its own getUserMedia video
    stream (no second permission prompt — browsers grant camera access per origin, not per call) and
    exposes capturePhoto(): string | null, which draws the current frame to an offscreen <canvas> and
    returns a base64 JPEG.
  • New useVoiceRecorder hook (ReactUi/src/hooks/): getUserMedia({ audio: true }) +
    MediaRecorder, exposing { isRecording, start, stop, error }; stop() resolves the recorded Blob as
    base64 via FileReader. First MediaRecorder/audio-getUserMedia usage in this codebase (confirmed
    via repo search — no prior audio capture existed), sibling to the existing camera hook rather than a
    generalized "media capture" abstraction, since the two have different lifecycles (continuous decode loop
    vs. start/stop-once recording).
  • New shared UnknownProductNameEntry component (ReactUi/src/components/) replaces the near-identical
    manual-<Input>-only fallback block duplicated today in PantryBarcodeScanner.tsx and
    ShoppingBarcodeScanDialog.tsx: the text input (unchanged), plus a " Foto" button (shown only when
    GetAiRecognitionCapabilitiesQuery().PhotoSupported) and a " Sprache" button (shown only when
    .SpeechSupported), each producing a suggestion that fills the same editable input the user already
    confirms/corrects before submitting — worth extracting now specifically because Phase 2 adds real
    behavior (camera capture, recording, two new network calls, a privacy notice) to what was one <Input>;
    copy-pasting that into both call sites would have tripled the new logic instead of sharing it.
  • A fixed, always-visible one-line notice directly above the two new buttons (not a one-time dismissible
    toast — see security pre-review point 3 below for why): "Foto/Aufnahme wird an den konfigurierten
    KI-Dienst gesendet."

Out of scope (per story)

No admin/settings UI for the AI config itself — this repo has no precedent anywhere for exposing
deployment-level external-service config in the frontend (SMTP/push VAPID are both env-only, confirmed by
grep), and the human's own answer asked for env/URL+key configurability, not a UI. No vendor selection UI.
No retry/fallback chain across multiple configured providers — one configured endpoint, full stop; an
operator who wants a specific provider's exact behavior configures that provider's URL directly.

**design** (`95_pantry_ai_product_recognition_phase2_design.md`) # Design: `#95` — Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2) **Architect note (2026-08-11).** Human resolved the story's own open question: use a common, self-hostable HTTP standard rather than a single hardcoded vendor SDK, with the endpoint, key, and model names all deployment-configurable, and Ollama specifically supported. That points at the **OpenAI-compatible API shape** — the de facto standard every major self-hosted inference server (Ollama, LocalAI, llama.cpp's server, vLLM) and every cloud vendor with an "OpenAI-compatible" mode already implements: `POST {base}/chat/completions` for vision (image content parts), `POST {base}/audio/transcriptions` for speech-to-text (Whisper-shaped multipart upload). One `HttpClient`, one base URL, one API key, two independently-configurable model names. ## Scope carried over from Phase 1 `#94`'s `ScanPantryProductBarcodeCommand` already has the exact mechanism the story's 4th AC asks for ("Das bestätigte Ergebnis wird wie in Phase 1 dauerhaft mit dem Barcode verknüpft gespeichert, nie zweimal gefragt"): when a barcode is unknown to both the pantry's own products (dedup by `(PantryId, Barcode)`) and Open Food Facts, the caller supplies `FallbackName` and a new `PantryProductEntity` permanently linked to that barcode is created. **Phase 2 does not touch that persistence path at all** — it only adds two new ways to *produce* the string the user was already typing by hand: a photo-recognition suggestion and a speech-transcription suggestion, both still shown to the user for confirmation/correction before the existing submit path runs. Same for Shopping's `ScanShoppingProductBarcodeCommand` (`#115` already shares Open Food Facts lookup with Pantry; this extends that sharing to the new AI lookups too). ## New deployment config — `AiSettings` ``` Ai:BaseUrl e.g. https://api.openai.com/v1, http://ollama-host:11434/v1, or a self-hosted Whisper-compatible server's URL. Empty = feature fully disabled. Ai:ApiKey Sent as `Authorization: Bearer {key}` when non-empty. Ollama ignores it; most cloud vendors require it. Ai:VisionModel e.g. "gpt-4o-mini", "llava", "moondream", "minicpm-v". Empty = photo recognition disabled even if BaseUrl is set. Ai:TranscriptionModel e.g. "whisper-1". Empty = speech recognition disabled even if BaseUrl is set. ``` Two independent `bool` flags, not one: `IsPhotoRecognitionConfigured` / `IsSpeechRecognitionConfigured`, each requiring `BaseUrl` **and** its own model name. This is a deliberate design response to the human's own "should work with local Ollama" requirement — Ollama's OpenAI-compatible layer serves `/chat/completions` (including vision models) but has no `/audio/transcriptions` endpoint at all. An operator running Ollama-only sets `VisionModel` and leaves `TranscriptionModel` blank; they get the photo option only, not a "Sprache" button that always 501s. Same env/appsettings/docker-compose wiring pattern as `Email`/`Push` (`.env.example` documents `AI_BASE_URL`/`AI_API_KEY`/`AI_VISION_MODEL`/ `AI_TRANSCRIPTION_MODEL`, mapped through `docker-compose.yml`'s `Ai__BaseUrl` etc.). No `ValidateOnStart()` — unlike `AppSettings.FrontendBaseUrl`, "all blank" is a valid, fully-supported "feature off" state, not a misconfiguration. ## New backend surface - `IAiProductRecognitionClient` (`CqsTodo/Ai/`), one `HttpClient`-backed implementation: - `TryRecognizeProductFromPhoto(byte[] jpegBytes, CancellationToken)` — one-shot chat completion, a fixed system prompt constraining the model to respond with a short grocery-product name and nothing else (no free-form chat, no injected user text — the only user-controlled input is the image itself), image sent as a `data:image/jpeg;base64,...` content part per the OpenAI vision message shape. - `TryTranscribeSpeech(byte[] audioBytes, string contentType, CancellationToken)` — multipart upload to `/audio/transcriptions` (`file`, `model`), returns the transcript text verbatim (the frontend dialog still shows it in an editable field — a mis-transcription is just a wrong prefill, not a data-integrity issue, same trust level Open Food Facts' suggestion already has). - Both: 30s timeout (vision/ASR inference is slower than Open Food Facts' 5s keyless lookup), both never throw out to the caller — any HTTP/timeout/malformed-response failure is logged and swallowed to `null`, matching `OpenFoodFactsClient`'s own "external outage never breaks the feature" convention. Whichever modality's suggestion comes back `null`, the user still has the plain manual-entry text field as the ultimate fallback — nothing about this feature can make entering a product name *harder* than Phase 1 already made it. - Three new public requests (own file each, standalone — reusable by both Pantry and Shopping, not nested under either feature folder): - `GetAiRecognitionCapabilitiesQuery() -> AiRecognitionCapabilitiesDto(bool PhotoSupported, bool SpeechSupported)` — lets the frontend hide a button that would always fail instead of discovering that by calling it and getting `null` back. - `RecognizeProductNameFromPhotoQuery(string ImageBase64) -> string?` - `RecognizeProductNameFromSpeechQuery(string AudioBase64, string ContentType) -> string?` - All three `WithAuthorization(_ => new AuthorizeIsCurrentUserAuthenticatedQuery())` — login-gated like push subscriptions/API keys, but deliberately *not* pantry/shopping-list-scoped, since recognizing "what's in this photo" needs no access to a specific list's data; the existing `AuthorizePantryAccessForCurrentUserQuery`/`AuthorizeShoppingListAccessForCurrentUserQuery` decorators still gate the actual `Scan*Command` that persists the confirmed name against a barcode. - Upload validation, shared with the existing avatar-upload path rather than re-implemented: extracted `UpdateUserAvatarCommandHandler.ProcessUpload`'s decode/validate core (base64 parse, size cap, format allowlist via `Image.Identify`, dimension cap, animated-frame rejection, `UnknownImageFormatException`/ `InvalidImageContentException` handling) into a shared `ImageUploadValidator.LoadAndValidate(base64, maxBytes, maxDimension)` returning a validated, decoded `Image`; each caller does its own encode/resize afterwards (avatar: crop-to-square 256px WebP; AI photo: cap-to-1024px-longest-side JPEG, no cropping — cropping a kitchen product photo could cut off the exact label text the model needs to read). Same two guards apply to the AI path as already apply to avatars: a decompression-bomb-shaped file is rejected before the full decode by dimension-checking by `Image.Identify` first, and the re-encode step means whatever EXIF/metadata (including GPS, if a phone attaches it) the original photo carried is never forwarded to the external AI endpoint — only re-encoded pixel data is. - Audio gets an equivalent but new, smaller validator (`AudioUploadValidator`) — no existing decode/re-encode path to extend, and no audio-processing library in this repo's dependency set, so this intentionally does *not* re-encode audio (there's nothing here to strip — `MediaRecorder`-produced WebM/Ogg has no EXIF-equivalent metadata risk). Enforces a size cap (10 MB — a few minutes of browser-recorded Opus easily fits, avoids a large-request DoS surface, similar reasoning to the photo's 2 MB cap) and a magic-byte content-type sniff against an allowlist (WebM/Ogg/WAV/MP3 signatures) so a client-forged `Content-Type` claim can't smuggle an arbitrary file past the size check into the outbound multipart request to the external endpoint. ## New frontend surface - New `usePhotoCapture` hook (`ReactUi/src/hooks/`), a sibling to `useBarcodeScanner` rather than an extension of it — found while implementing that by the time a user reaches the "unknown product" fallback, `useBarcodeScanner`'s own successful-decode callback has already called `controls.stop()`, so there's no still-open stream left to snapshot from (the design's original plan to extend that hook didn't hold up against its actual lifecycle). `usePhotoCapture` opens its own `getUserMedia` video stream (no second permission *prompt* — browsers grant camera access per origin, not per call) and exposes `capturePhoto(): string | null`, which draws the current frame to an offscreen `<canvas>` and returns a base64 JPEG. - New `useVoiceRecorder` hook (`ReactUi/src/hooks/`): `getUserMedia({ audio: true })` + `MediaRecorder`, exposing `{ isRecording, start, stop, error }`; `stop()` resolves the recorded Blob as base64 via `FileReader`. First `MediaRecorder`/audio-`getUserMedia` usage in this codebase (confirmed via repo search — no prior audio capture existed), sibling to the existing camera hook rather than a generalized "media capture" abstraction, since the two have different lifecycles (continuous decode loop vs. start/stop-once recording). - New shared `UnknownProductNameEntry` component (`ReactUi/src/components/`) replaces the near-identical manual-`<Input>`-only fallback block duplicated today in `PantryBarcodeScanner.tsx` and `ShoppingBarcodeScanDialog.tsx`: the text input (unchanged), plus a " Foto" button (shown only when `GetAiRecognitionCapabilitiesQuery().PhotoSupported`) and a " Sprache" button (shown only when `.SpeechSupported`), each producing a suggestion that fills the same editable input the user already confirms/corrects before submitting — worth extracting now specifically because Phase 2 adds real behavior (camera capture, recording, two new network calls, a privacy notice) to what was one `<Input>`; copy-pasting that into both call sites would have tripled the new logic instead of sharing it. - A fixed, always-visible one-line notice directly above the two new buttons (not a one-time dismissible toast — see security pre-review point 3 below for why): *"Foto/Aufnahme wird an den konfigurierten KI-Dienst gesendet."* ## Out of scope (per story) No admin/settings UI for the AI config itself — this repo has no precedent anywhere for exposing deployment-level external-service config in the frontend (SMTP/push VAPID are both env-only, confirmed by grep), and the human's own answer asked for env/URL+key configurability, not a UI. No vendor selection UI. No retry/fallback chain across multiple configured providers — one configured endpoint, full stop; an operator who wants a specific provider's exact behavior configures that provider's URL directly.
lena commented 2026-08-18 13:13:57 +02:00 (Migrated from git.butzei.de)

security_prereview (95_pantry_ai_product_recognition_phase2_security_prereview.md)

Security Pre-Review: #95 — Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2)

Reviewed before implementation, per ai/roles/00_team_overview.md's feature cycle (design → security
pre-review → implementation) and this story's own explicit "Sicherheits-Vorprüfung ist für diese Story
verpflichtend, nicht optional" clause. This is the first feature in the app that sends user-originated
photos/audio to an external third party, so the story's own required questions are answered directly
below before anything else.

Story's own required questions

Welcher Anbieter? None fixed — the human's resolution deliberately made this a deployment choice, not
a build-time one (Ai:BaseUrl/Ai:VisionModel/Ai:TranscriptionModel, blank = feature off). Whoever runs
this app instance chooses their own endpoint: a self-hosted Ollama (data never leaves their own
infrastructure), a self-hosted Whisper-compatible transcription server, or a cloud vendor with an
OpenAI-compatible API.

Welche Daten verlassen die App? A JPEG-re-encoded photo (stripped of any original metadata/EXIF/GPS —
see point 2) when "Foto" is used, or a WebM/Ogg/WAV/MP3 audio clip when "Sprache" is used. Nothing else —
no session cookie, no user id, no pantry/list identifiers are included in either outbound request (see
point 4).

Werden sie beim Anbieter gespeichert/für Training verwendet? Unknowable by this app at build time
— it depends entirely on whichever endpoint the operator configures, and is exactly the kind of fact this
app cannot verify or enforce technically. This is answered by disclosure, not by code: the in-app notice
(point 3) tells the user that their photo/recording leaves the app to "the configured AI service" — it
is the deploying operator's responsibility (same as choosing an SMTP relay or a Seq log sink today) to
pick an endpoint whose data-handling policy they're comfortable with, and to document that for their own
users if it differs from "processed only, never retained." This app does not claim a specific vendor's
policy, since it doesn't know which vendor will be configured.

Braucht es vor der ersten Nutzung einen Zustimmungs-/Datenschutzhinweis? Yes — required, not
optional.
See point 3 for why it must be always-visible rather than a one-time dismissible dialog.

1. Blast radius of a "wrong" AI response

Risk: A malicious or malfunctioning configured endpoint returns attacker-controlled text as the
"recognized" product name or transcript.
Mitigation: The returned string is never trusted further than Phase 1 already trusts Open Food
Facts' response (see 94_..._security_prereview.md point 5) or a user's own typed input: it only ever
prefills the same editable text field the user must still confirm before submitting, and on submission it
passes through the exact same PantryProductName/ShoppingProductName Vogen validator (length cap, no
special-casing) as manual entry. It is rendered only as escaped React text, never as markup/HTML. It can
never reach a SQL query, a shell command, or any privileged action — the only thing "downstream" of this
string is a plain varchar column.

2. Photo upload — decompression bombs, format spoofing, metadata leakage

Risk: Same class of risk as the existing avatar upload (UpdateUserAvatarCommandHandler): an
oversized or maliciously-crafted image could exhaust server memory during decode, or forward
identifying metadata (GPS EXIF from a phone photo of the kitchen) to the external AI endpoint.
Mitigation: Reuses the avatar path's exact validation core (extracted into
ImageUploadValidator.LoadAndValidate): base64-decode with a byte-length cap before any image decode,
Image.Identify (header-only, cheap) checked against dimensions and an allowlist of decodable formats
before the full Image.Load, and animated multi-frame images rejected — all before the expensive full
decode ever runs, closing the same "many-large-frames-in-a-small-file" bomb vector the avatar review
already closed. The photo is then always re-encoded to plain JPEG before being sent onward — the original
file's bytes (and anything embedded in them) are discarded entirely; only decoded pixel data survives the
round-trip, so no EXIF/GPS/embedded-metadata of any kind reaches the external endpoint even if the source
photo carried it.

3. Audio upload — format spoofing, size DoS

Risk: An arbitrarily large or mislabeled file forwarded to an external endpoint as "audio" (resource
exhaustion, or smuggling an unrelated file type past a naive Content-Type-trusting check).
Mitigation: Size cap enforced on the decoded byte length before any further processing (10 MB — well
above a realistic few-minutes-long browser voice recording). Content type is verified by sniffing the
first bytes against known container signatures (WebM/EBML, OggS, RIFF/WAVE, MP3 frame sync/ID3) rather
than trusting the client-supplied MIME string, so a forged Content-Type header can't bypass the
allowlist. No re-encoding is done (no audio library in this repo's dependencies, and browser-recorded
WebM/Opus carries no EXIF-equivalent identifying metadata the way a phone photo does) — the size cap and
signature check are the full mitigation here, proportionate to the actual risk shape.

4. Outbound request — SSRF surface, credential handling

Risk: Unlike Open Food Facts' fixed, code-constant base URL, Ai:BaseUrl is operator-configurable —
but a config value the operator sets themselves (like Email:Smtp:Host or Push:Vapid:* already are) is
a fundamentally different trust boundary than user-supplied input. No part of the base URL/host is ever
derived from a request the frontend sends — a logged-in user cannot redirect the outbound call anywhere,
regardless of what they submit as ImageBase64/AudioBase64. This is the same trust model this codebase
already applies to the DB connection string and SMTP host: an operator with config-file/environment access
is not an untrusted party in this app's threat model.
Mitigation: Ai:ApiKey, when set, is sent only as an Authorization: Bearer header to the
configured Ai:BaseUrl — never logged (handlers log request/response shapes — "recognition
succeeded/failed" — never the key, the image bytes, the audio bytes, or the returned text verbatim beyond
what's already user-visible in the UI it prefills). No session cookie, auth token, user id, or any other
app-internal identifier is included in the outbound request to the AI endpoint — the request carries only
the media bytes and a fixed system prompt, so a compromised or malicious configured endpoint learns
nothing about the app's users beyond the content of whatever they chose to photograph/say. 30s timeout +
try/catch around the whole call, same "external outage can never break or hang the feature" convention as
OpenFoodFactsClient.

5. Prompt injection via image/audio content

Risk: A crafted image (text overlay) or spoken phrase could attempt to make the underlying model
ignore its system prompt and return something other than a plausible product name — e.g. an
attacker-controlled long string.
Mitigation: Bounded blast radius, not prevented outright (prompt injection against a
third-party-hosted model isn't something this app's own code can fully close, and doesn't need to be —
see point 1): the model's entire "authority" is a single string returned into a user-editable text field
that the user reviews before it becomes a PantryProductName/ShoppingProductName, capped by that value
object's existing length validation exactly like every other name in this app, manual or suggested. The
worst case is a nonsense or overlong-and-truncated suggestion the user simply retypes — functionally
identical to Open Food Facts returning a bad name today, already an accepted, non-blocking risk shape in
this codebase (94_..._security_prereview.md point 5).

Risk: A dismiss-once "I understand" dialog (like a typical cookie banner) would stop being accurate
information the moment an operator changes their configured endpoint — a user who dismissed it once
under "data stays local (Ollama)" wouldn't be re-notified if the operator later switches to a cloud vendor.
Mitigation: The notice is a fixed, always-rendered line directly above the Foto/Sprache buttons in
UnknownProductNameEntry, not a dismissible interstitial — it costs nothing to show every time (this flow
is already a rare, occasional "unknown barcode" path, not a hot path where repetition would be
disruptive), and never goes stale relative to whatever endpoint happens to be configured right now.

Ergebnis

No blocking findings. Point 4's SSRF concern is resolved by trust-boundary reasoning (operator config,
not user input) rather than technical prevention, consistent with how this codebase already treats every
other operator-supplied external-service endpoint. Approved for implementation as designed.

**security_prereview** (`95_pantry_ai_product_recognition_phase2_security_prereview.md`) # Security Pre-Review: `#95` — Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2) Reviewed before implementation, per `ai/roles/00_team_overview.md`'s feature cycle (design → security pre-review → implementation) and this story's own explicit "Sicherheits-Vorprüfung ist für diese Story verpflichtend, nicht optional" clause. This is the first feature in the app that sends user-originated photos/audio to an external third party, so the story's own required questions are answered directly below before anything else. ## Story's own required questions **Welcher Anbieter?** None fixed — the human's resolution deliberately made this a deployment choice, not a build-time one (`Ai:BaseUrl`/`Ai:VisionModel`/`Ai:TranscriptionModel`, blank = feature off). Whoever runs this app instance chooses their own endpoint: a self-hosted Ollama (data never leaves their own infrastructure), a self-hosted Whisper-compatible transcription server, or a cloud vendor with an OpenAI-compatible API. **Welche Daten verlassen die App?** A JPEG-re-encoded photo (stripped of any original metadata/EXIF/GPS — see point 2) when "Foto" is used, or a WebM/Ogg/WAV/MP3 audio clip when "Sprache" is used. Nothing else — no session cookie, no user id, no pantry/list identifiers are included in either outbound request (see point 4). **Werden sie beim Anbieter gespeichert/für Training verwendet?** **Unknowable by this app at build time** — it depends entirely on whichever endpoint the operator configures, and is exactly the kind of fact this app cannot verify or enforce technically. This is answered by disclosure, not by code: the in-app notice (point 3) tells the *user* that their photo/recording leaves the app to "the configured AI service" — it is the *deploying operator's* responsibility (same as choosing an SMTP relay or a Seq log sink today) to pick an endpoint whose data-handling policy they're comfortable with, and to document that for their own users if it differs from "processed only, never retained." This app does not claim a specific vendor's policy, since it doesn't know which vendor will be configured. **Braucht es vor der ersten Nutzung einen Zustimmungs-/Datenschutzhinweis?** **Yes — required, not optional.** See point 3 for why it must be always-visible rather than a one-time dismissible dialog. ## 1. Blast radius of a "wrong" AI response **Risk:** A malicious or malfunctioning configured endpoint returns attacker-controlled text as the "recognized" product name or transcript. **Mitigation:** The returned string is never trusted further than Phase 1 already trusts Open Food Facts' response (see `94_..._security_prereview.md` point 5) or a user's own typed input: it only ever prefills the same editable text field the user must still confirm before submitting, and on submission it passes through the exact same `PantryProductName`/`ShoppingProductName` Vogen validator (length cap, no special-casing) as manual entry. It is rendered only as escaped React text, never as markup/HTML. It can never reach a SQL query, a shell command, or any privileged action — the only thing "downstream" of this string is a plain varchar column. ## 2. Photo upload — decompression bombs, format spoofing, metadata leakage **Risk:** Same class of risk as the existing avatar upload (`UpdateUserAvatarCommandHandler`): an oversized or maliciously-crafted image could exhaust server memory during decode, or forward identifying metadata (GPS EXIF from a phone photo of the kitchen) to the external AI endpoint. **Mitigation:** Reuses the avatar path's exact validation core (extracted into `ImageUploadValidator.LoadAndValidate`): base64-decode with a byte-length cap *before* any image decode, `Image.Identify` (header-only, cheap) checked against dimensions and an allowlist of decodable formats before the full `Image.Load`, and animated multi-frame images rejected — all before the expensive full decode ever runs, closing the same "many-large-frames-in-a-small-file" bomb vector the avatar review already closed. The photo is then always re-encoded to plain JPEG before being sent onward — the original file's bytes (and anything embedded in them) are discarded entirely; only decoded pixel data survives the round-trip, so no EXIF/GPS/embedded-metadata of any kind reaches the external endpoint even if the source photo carried it. ## 3. Audio upload — format spoofing, size DoS **Risk:** An arbitrarily large or mislabeled file forwarded to an external endpoint as "audio" (resource exhaustion, or smuggling an unrelated file type past a naive `Content-Type`-trusting check). **Mitigation:** Size cap enforced on the decoded byte length before any further processing (10 MB — well above a realistic few-minutes-long browser voice recording). Content type is verified by sniffing the first bytes against known container signatures (WebM/EBML, OggS, RIFF/WAVE, MP3 frame sync/ID3) rather than trusting the client-supplied MIME string, so a forged `Content-Type` header can't bypass the allowlist. No re-encoding is done (no audio library in this repo's dependencies, and browser-recorded WebM/Opus carries no EXIF-equivalent identifying metadata the way a phone photo does) — the size cap and signature check are the full mitigation here, proportionate to the actual risk shape. ## 4. Outbound request — SSRF surface, credential handling **Risk:** Unlike Open Food Facts' fixed, code-constant base URL, `Ai:BaseUrl` is *operator*-configurable — but a config value the operator sets themselves (like `Email:Smtp:Host` or `Push:Vapid:*` already are) is a fundamentally different trust boundary than *user*-supplied input. No part of the base URL/host is ever derived from a request the frontend sends — a logged-in user cannot redirect the outbound call anywhere, regardless of what they submit as `ImageBase64`/`AudioBase64`. This is the same trust model this codebase already applies to the DB connection string and SMTP host: an operator with config-file/environment access is not an untrusted party in this app's threat model. **Mitigation:** `Ai:ApiKey`, when set, is sent only as an `Authorization: Bearer` header to the configured `Ai:BaseUrl` — never logged (handlers log request/response *shapes* — "recognition succeeded/failed" — never the key, the image bytes, the audio bytes, or the returned text verbatim beyond what's already user-visible in the UI it prefills). No session cookie, auth token, user id, or any other app-internal identifier is included in the outbound request to the AI endpoint — the request carries only the media bytes and a fixed system prompt, so a compromised or malicious configured endpoint learns nothing about the app's users beyond the content of whatever they chose to photograph/say. 30s timeout + try/catch around the whole call, same "external outage can never break or hang the feature" convention as `OpenFoodFactsClient`. ## 5. Prompt injection via image/audio content **Risk:** A crafted image (text overlay) or spoken phrase could attempt to make the underlying model ignore its system prompt and return something other than a plausible product name — e.g. an attacker-controlled long string. **Mitigation:** Bounded blast radius, not prevented outright (prompt injection against a third-party-hosted model isn't something this app's own code can fully close, and doesn't need to be — see point 1): the model's entire "authority" is a single string returned into a user-editable text field that the user reviews before it becomes a `PantryProductName`/`ShoppingProductName`, capped by that value object's existing length validation exactly like every other name in this app, manual or suggested. The worst case is a nonsense or overlong-and-truncated suggestion the user simply retypes — functionally identical to Open Food Facts returning a bad name today, already an accepted, non-blocking risk shape in this codebase (`94_..._security_prereview.md` point 5). ## 6. Consent notice — why always-visible, not a one-time dialog **Risk:** A dismiss-once "I understand" dialog (like a typical cookie banner) would stop being accurate information the moment an operator changes their configured endpoint — a user who dismissed it once under "data stays local (Ollama)" wouldn't be re-notified if the operator later switches to a cloud vendor. **Mitigation:** The notice is a fixed, always-rendered line directly above the Foto/Sprache buttons in `UnknownProductNameEntry`, not a dismissible interstitial — it costs nothing to show every time (this flow is already a rare, occasional "unknown barcode" path, not a hot path where repetition would be disruptive), and never goes stale relative to whatever endpoint happens to be configured right now. ## Ergebnis No blocking findings. Point 4's SSRF concern is resolved by trust-boundary reasoning (operator config, not user input) rather than technical prevention, consistent with how this codebase already treats every other operator-supplied external-service endpoint. Approved for implementation as designed.
lena commented 2026-08-18 13:13:57 +02:00 (Migrated from git.butzei.de)

security_final (95_pantry_ai_product_recognition_phase2_security_final.md)

Security Final Review: #95 — Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2)

Performed after implementation, per ai/roles/00_team_overview.md's feature cycle. A dedicated
security-review subagent traced authorization wiring, upload validation, the outbound HTTP call, and
data flow for the returned AI suggestion end to end. No high-confidence findings.

Findings

None. Verified specifically:

  • Authorization: all three new endpoints (GetAiRecognitionCapabilitiesQueryHandler,
    RecognizeProductNameFromPhotoQueryHandler, RecognizeProductNameFromSpeechQueryHandler) wire
    WithAuthorization(_ => new AuthorizeIsCurrentUserAuthenticatedQuery()), identical to the
    established bare-login-gate pattern (GetPushVapidPublicKeyQueryHandler and 30+ siblings). None
    touch per-user/per-list data, so there is no cross-user data exposure.
  • Upload validation: ImageUploadValidator.LoadAndValidate re-decodes bytes independent of any
    claimed content type, enforces the size cap, checks the actual decoded format against an allowlist,
    bounds pixel dimensions, and rejects animated images before the full decode. AudioUploadValidator
    does real magic-byte container sniffing (WebM/Ogg/WAV/MP3), never trusting a client-supplied MIME
    string.
  • Avatar handler regression check: diffed UpdateUserAvatarCommandHandler.cs before/after the
    ImageUploadValidator extraction — the exact same size cap (2 MB), format allowlist,
    MaxSourcePixelDimension (4096), and animated-frame rejection survived verbatim. No regression from
    the refactor.
  • Outbound HTTP call / SSRF: Ai:BaseUrl is read only from server configuration at DI-registration
    time; no user-supplied value ever reaches HttpClient.BaseAddress or influences protocol/host.
  • Prompt injection / data flow to sinks: the vision system prompt has no user-controlled text
    concatenated into it — only image bytes are attacker-influenced. The AI's returned suggestion flows
    into a React-controlled <Input> (no dangerouslySetInnerHTML anywhere in the changed files) and
    still passes through the same Vogen validators as manually-typed input before persistence.
  • Data exposure / logging: AiProductRecognitionClient logs only exception objects on failure,
    never request/response bodies. The AI API key is only ever added as an Authorization: Bearer header
    on the dedicated outbound HttpClient, never logged or returned to the frontend —
    AiRecognitionCapabilitiesDto exposes only two booleans, not the base URL or key.

Result

No blocking findings. Approved as implemented. See
95_pantry_ai_product_recognition_phase2_security_prereview.md for the pre-implementation review this
follows up on.

**security_final** (`95_pantry_ai_product_recognition_phase2_security_final.md`) # Security Final Review: `#95` — Vorratsschrank KI-Erkennung unbekannter Produkte per Foto/Sprache (Phase 2) Performed after implementation, per `ai/roles/00_team_overview.md`'s feature cycle. A dedicated security-review subagent traced authorization wiring, upload validation, the outbound HTTP call, and data flow for the returned AI suggestion end to end. No high-confidence findings. ## Findings None. Verified specifically: - **Authorization**: all three new endpoints (`GetAiRecognitionCapabilitiesQueryHandler`, `RecognizeProductNameFromPhotoQueryHandler`, `RecognizeProductNameFromSpeechQueryHandler`) wire `WithAuthorization(_ => new AuthorizeIsCurrentUserAuthenticatedQuery())`, identical to the established bare-login-gate pattern (`GetPushVapidPublicKeyQueryHandler` and 30+ siblings). None touch per-user/per-list data, so there is no cross-user data exposure. - **Upload validation**: `ImageUploadValidator.LoadAndValidate` re-decodes bytes independent of any claimed content type, enforces the size cap, checks the actual decoded format against an allowlist, bounds pixel dimensions, and rejects animated images before the full decode. `AudioUploadValidator` does real magic-byte container sniffing (WebM/Ogg/WAV/MP3), never trusting a client-supplied MIME string. - **Avatar handler regression check**: diffed `UpdateUserAvatarCommandHandler.cs` before/after the `ImageUploadValidator` extraction — the exact same size cap (2 MB), format allowlist, `MaxSourcePixelDimension` (4096), and animated-frame rejection survived verbatim. No regression from the refactor. - **Outbound HTTP call / SSRF**: `Ai:BaseUrl` is read only from server configuration at DI-registration time; no user-supplied value ever reaches `HttpClient.BaseAddress` or influences protocol/host. - **Prompt injection / data flow to sinks**: the vision system prompt has no user-controlled text concatenated into it — only image bytes are attacker-influenced. The AI's returned suggestion flows into a React-controlled `<Input>` (no `dangerouslySetInnerHTML` anywhere in the changed files) and still passes through the same Vogen validators as manually-typed input before persistence. - **Data exposure / logging**: `AiProductRecognitionClient` logs only exception objects on failure, never request/response bodies. The AI API key is only ever added as an `Authorization: Bearer` header on the dedicated outbound `HttpClient`, never logged or returned to the frontend — `AiRecognitionCapabilitiesDto` exposes only two booleans, not the base URL or key. ## Result No blocking findings. Approved as implemented. See `95_pantry_ai_product_recognition_phase2_security_prereview.md` for the pre-implementation review this follows up on.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
robert/todo#95
No description provided.