# DeepSeek V4.1 Flash staging evaluation

11 September 2026 · Round two · Direct LLM review · Routing awaiting approval

**Recommendation: do not make a blanket replacement.** Three services are scoped candidates, five need more evidence, and 33 remain on hold. All 243 available test cases have a final result and a written review. The requested historical coverage is still incomplete; fresh comparisons and repeated batches are labeled explicitly.

| Measure | Result |
|---|---|
| Selected CSV rows / evaluated services | 29 / 41 |
| DeepSeek final responses / failed cases | 242 / 1 |
| Historical reconstructions / fresh pairs | 159 / 84 |
| Fresh existing-primary responses | 84 / 84 |
| Per-case wall time: median / p95 | 19.7 s / 294.7 s |
| Known returned round-two DeepSeek cost | $0.5259; incomplete spend |

The one failed case is AI director `r29-ai-director-4828d071-78ba-4a27-83c5-a00030154df7-round2`: five connection-error attempts, 1060.3 s elapsed. No result from an existing model replaces it in these numbers. The earlier clinic photoshoot Gemini 503 is archived; its retry returned successfully and is included in the direct comparison.

**How to read decisions:** Candidate = reasonable for narrow approval consideration; Limited candidate = observed usefulness but insufficient or mixed coverage; Hold = observed material failure and/or evidence gaps prevent a migration recommendation. These are direct judgments, not automated grades.

## Scoped candidates

- **Storyboard style patch:** Both models make the requested narrow style edits and refuse the two changes that would alter video shape. DeepSeek gives useful compiled descriptions; no broad quality advantage is established. Evidence: 0 historical reconstructions, 6 fresh pairs.
- **Attachment descriptions:** The descriptions capture actual poster text, layout and screenshot context more usefully than the saved broad scene labels. Small descriptive inaccuracies remain, but no material invented business offering appeared in this sample. Evidence: 6 historical reconstructions, 0 fresh pairs.
- **Website theme suggestions:** All six sets offer a credible minimal, bold and professional direction. DeepSeek often separates the bold palette more clearly than the saved sets. These are design proposals, not verified brand facts or rendered/contrast-tested themes. Evidence: 6 historical reconstructions, 0 fresh pairs.

## Service decisions

| Row | Service | Decision | Historical / fresh | DeepSeek thinking | Main finding |
|---|---|---|---|---|---|
| 2 | Business text analysis | Hold | 0 / 6 | disabled | Fresh pairs use identical scraped text, but extraction must not turn inference into business facts. WallFusion gains a money-back guarantee and handcrafted production that the supplied text does not establish; logo fallback handling also needs repair. |
| 3 | Website visual identity | Hold | 6 / 0 | disabled | DeepSeek notices the obstructed pages more clearly than historical outputs. However, the browser-check case still fills brand palette and layout fields using generic interstitial UI despite saying no genuine brand identity is available. Downstream consumers cannot safely rely on the note alone. |
| 4 | Scraped-photo selection | Hold | 0 / 6 | disabled | This is a metadata-only task. DeepSeek correctly filters many obvious logos/icons and selects sensible product indices, but its own BIBA reasons explicitly accept navigation/collection banners despite the rejection rule. 9Blings URLs also expose a capture issue: width/height=1 despite large declared dimensions. |
| 5 | Business photo analysis | Hold | 6 / 0 | disabled | DeepSeek produces more specific visual descriptions and usable concepts, but labels a kurta photo as food, invents QR booking, opening status and handcraft/fabric claims. These errors would propagate into downstream marketing. |
| 6 | Fashion classification | Limited candidate | 6 / 0 | disabled | Demographic labels match the saved outputs. Western dominance for SMARTLOOK is defensible despite one kurta-like shirt. Captor incorrectly discounts accessories, although this service explicitly includes accessory businesses; do not use its confidence to suppress those businesses without calibration. |
| 6 | Quick photo triage | Limited candidate | 6 / 0 | disabled | Both models accept six usable references and identify their broad subjects. No rejection examples were sampled, so false acceptance of logos, trackers or irrelevant media remains untested. Reference scores differ enough to affect selection thresholds. |
| 7 | Brand intelligence | Hold | 0 / 6 | high | Fresh identical-input comparisons show worthwhile strategy synthesis. Iconic Vision gains single-item ordering, hand finishing, in-house shipping and shorter lead times without evidence. Aainu gains supervised/on-site management. These claims would propagate into downstream content. |
| 7 | Brand visual system | Hold | 6 / 0 | high | Several useful art directions, especially Iconic Vision and 10TH COSMOS. However, Max Lab explicitly downgrades a palette labelled real, and Aainu introduces contradictory white-space targets and bans colours present in natural photography. This is not yet a consistently usable short visual brief. |
| 8 | Monthly calendar strategy | Hold | 6 / 0 | high | The plans broadly follow the 22-date window, evergreen emphasis and varied subjects. However, new arrivals, actual production procedures and business-specific advice are asserted without supporting facts. Some historical outputs are also poorly grounded. Festival quality cannot be compared because the historical candidate list was not retained. |
| 9 | Calendar variants | Hold | 6 / 0 | high | All six requests are for Max Lab and reconstruct initial generation without prior variants or regeneration instructions. DeepSeek repeatedly invents a link-in-bio booking route, strengthens accreditation into blanket trust claims and asserts local fever trends without evidence. The outputs do not establish safe regeneration behavior. |
| 10 | Caption shortening | Limited candidate | 0 / 6 | low | All six retain the main message, CTA and hashtags while removing repeated blocks. This is a weak compression challenge made from one lab’s repeated captions, not six real shortening requests; diverse nonrepetitive and multilingual captions remain untested. |
| 10 | Copy after photo changes | Hold | 0 / 6 | low | Six fresh tests all replace references with the same Max Lab logo. This is narrow coverage of an image-change task. Most copies remain on theme; the monsoon version becomes repetitive and turns generic pathology services into a seasonal recovery check-up. The baseline is clearer in that case. |
| 11 | Calendar refresh | Hold | 5 / 0 | high | Five batches are one independent staging run, not five runs. The source mixes a dot business name, fashion category, software-security claims and unrelated winery/preschool/fashion references. DeepSeek uses the assigned descriptions more literally and often preserves copy, but produces incoherent combinations and newly invents a seasonal phishing claim. Missing previous facts also prevent reliable removed-service evaluation. |
| 12 | Batch video ideas | Hold | 6 / 0 | high | The six complete 20-idea batches are readable and often actionable, but DeepSeek invents timings, founder history and customer outcomes, repeats angles, and sometimes requests an on-camera host while presenter is none. Source fact contamination and incomplete historical photo ordering also need correction before migration. |
| 12 | Onboarding video idea | Hold | 6 / 0 | high | The supplied hero descriptions and current profiles sometimes differ materially from saved outputs. DeepSeek produces focused 20-second concepts but can add unsupported product properties or review content. Resolve replay fidelity and these claims before replacing. |
| 12 | Trending video ideas | Hold | 0 / 6 | high | Different formats are recognizable and varied, but the restaurant gets fictitious regular-customer behavior and a hundred hot-delivered boxes, the shop gets invented opening times/process durations, and the clinic gets an invented four-second vaccination and dawn opening. Creativity does not make these publishable. |
| 14 | Storyboard compilation | Hold | 6 / 0 | high | DeepSeek produces coherent beat structures, but adds real-world proof claims, omits required asset references and exceeds spoken word budgets in several beats. The source prompt itself encourages unverified proof/caveats and has conflicting asset-description rules, so fixing only the provider will not solve this. |
| 16 | Storyboard style patch | Candidate | 0 / 6 | low | Both models make the requested narrow style edits and refuse the two changes that would alter video shape. DeepSeek gives useful compiled descriptions; no broad quality advantage is established. |
| 18 | Image prompt enhancement | Hold | 6 / 0 | high | The model can produce useful detailed prompts, but it converts a sitting-pose request into trousers on a hanger or flat surface with no person. Other cases introduce material guesses and copy claims. For an enhancer, preserving the requested transformation is essential. |
| 19 | Product pack briefs | Hold | 4 / 0 | high | Four outputs cover only two product plans. DeepSeek removes some unsupported material and dimension claims in the saved barbell plan, but the Flipkart prompts lose concrete camera distances and occasionally request unseen wear. New angles, scaling and product identity need generated-image evaluation before adoption. |
| 19 | Product pack refinement | Limited candidate | 0 / 6 | high | All six fresh tests apply the same softer-light edit to shots from only two saved plans. The requested lighting change is handled without moving the camera or changing required white backgrounds. Useful evidence for this narrow edit, not for arbitrary refinements. |
| 20 | Photoshoot ideas | Hold | 0 / 6 | high | Fresh tests show the model can create coherent outfit settings and preserve common garment details. It invents an AUSPORA fashion collection and turns a clinic into an ethnicwear set, with an unverified Ayushman Bharat story. The clinic Gemini retry stays closer to a healthcare theme, but also invents WHO posters, staff identities and an event. Both require better grounding; the original 503 is archived separately. |
| 22 | Attachment descriptions | Candidate | 6 / 0 | disabled | The descriptions capture actual poster text, layout and screenshot context more usefully than the saved broad scene labels. Small descriptive inaccuracies remain, but no material invented business offering appeared in this sample. |
| 22 | Carousel planning | Hold | 0 / 6 | disabled | Fresh image pairs expose unsupported cotton/breathability/screen-printing claims, invented KPD garments, a fusion-menu story and WHO references absent from the source. Several existing outputs also overclaim; neither should be treated as ground truth. |
| 23 | Native website generation | Hold | 6 / 0 | high | All six drafts prepare and validate after application recovery, with 7–19 warnings per draft. Copy is generally more concrete and less inflated than the saved sites, but unsupported product properties, stock/response promises and unsuitable media remain. Historical asset subsets differ; layout, accessibility and live interactions were not visually evaluated. Long retry-inclusive completion times also need attention. |
| 24 | Native website editing | Hold | 6 / 0 | low | Three operation plans apply and prepare successfully; three requests return clarification. The rename is broadly sensible, but the address edit updates a structured contact field while old-location editorial copy remains. The maps flow leaves an earlier requested address correction unresolved. Historical finalized documents do not reliably expose the original provider operation result. |
| 25 | Website builder prompt | Hold | 6 / 0 | disabled | All six preserve the main theme colours and image-URL rules. Missing historical Brand Positioning fields limit the comparison, yet DeepSeek independently instructs builders to add invented services, hours, phone numbers or fake form success. These are unacceptable default behaviors even with sparse inputs. |
| 25 | Website theme suggestions | Candidate | 6 / 0 | disabled | All six sets offer a credible minimal, bold and professional direction. DeepSeek often separates the bold palette more clearly than the saved sets. These are design proposals, not verified brand facts or rendered/contrast-tested themes. |
| 27 | Click-to-WhatsApp ad copy | Hold | 0 / 6 | low | The messages are generally usable, but DeepSeek asserts online-only/pan-India service when the profile merely lacks a location. That can mislead owners about targeting rationale. Sparse profiles need a safer uncertainty policy. |
| 27 | Engagement ad copy | Hold | 0 / 6 | low | DeepSeek more consistently asks a natural question suited to comments, while several existing outputs read like visit-the-store ads. However, two DeepSeek rationales invent pan-India shipping, which outweighs the creative advantage for unrestricted rollout. |
| 28 | Google post polishing | Hold | 0 / 6 | low | Text preservation is generally good, but KPD is labelled Polish without evidence. Five of the six drafts are only names or near-empty fragments, so this is weak evidence for real merchant-draft polishing. |
| 28 | Google post writing | Hold | 0 / 6 | low | The visual descriptions are generally accurate, but a profile/photo conflict produces unusable hybrid copy, the storefront case adds unsupported newness and questionable OCR, and the restaurant case invents signature status. Both models need stronger handling of ambiguous inputs. |
| 29 | AI video editing director | Hold | 6 / 0 | high | Five outputs and one exhausted connection failure. Returned directions invent face/product positions without a frame, and some timing conflicts are visible directly in the output. This service needs complete director baselines, frames and render validation. |
| 31 | Avatar gender labels | Limited candidate | 6 / 0 | high | All six prompts explicitly describe a man and both models return male. This is straightforward extraction, with no observed disagreement; the sample says nothing about female or ambiguous prompts. |
| 34 | Garment analysis | Hold | 6 / 0 | low | The main garment readings are mostly convincing. A visible background shirt is included as outerwear in one case but ignored in similar neighboring images; this could add unwanted clothing to generation. The red dress disagreement is plausibly a historical label problem, not evidence DeepSeek is wrong. No saree cases were sampled. |
| 34 | Garment prompt writing | Hold | 6 / 0 | low | Framing and background instructions are consistent across six outputs, and user no-pants overrides are preserved. One output drops the explicit face-reference instruction. All cases refer to an additional face image but the captured requests contain only the garment image, limiting identity evaluation. |
| 35 | Product analysis | Hold | 6 / 0 | low | DeepSeek supplies useful leaf labels, including barbell where the saved label says dumbbell, and specific shot concepts. But it proposes a wooden-door installation for what appears to be glass-door hardware and invents an unseen jacket back. Classification alone is more promising than the combined shot-idea output. |
| 35 | Product prompt writing | Hold | 6 / 0 | low | The output-key conflict and HEIC transport are fixed in the eval adapter. Direct quality review still finds invented fabric properties, an insufficient no-face instruction and a generic handheld treatment of a jacket. Formatting success is not enough to approve this service. |
| 36 | Avatar promotional script | Hold | 6 / 0 | low | The copy often sounds natural, and visible product text is used well. However, it fabricates personal/customer experiences and stronger exclusivity than supplied. Historical scripts sometimes include user edits and mismatched current business context, so this is not a clean model-only win-rate comparison. |
| 37 | Voice script processing | Hold | 6 / 0 | low | The preservation instruction improves the previous failure but does not fix it: the English production brief is still translated to Hinglish despite an explicit instruction to preserve English. This service should not migrate yet. |
| 38 | Sora prompt enhancement | Hold | 6 / 0 | high | All six return usable text structures. The tuition script is far too long for ten seconds, the five-second tee ad remains unspeakably dense, and the Pichola request loses its continuous full-length framing. The Mediterranean ad adds a pure-linen claim and changes the actor’s attire despite preservation instructions. |

## Thinking and model order

Disabled was requested for 60 Flash-Lite cases, low for 84 cases, and high for 99 cases. Existing OpenAI-first low paths retain low on DeepSeek even where the CSV lists a high Gemini fallback. Medium maps to high approximately; default/automatic non-Lite calls were tested at high. Photoshoot ideas used high against the existing image model’s default. The providers’ exposed reasoning-token counts are available per case, but neither these labels nor token counts establish equal internal computation. Existing fallback settings remain unchanged.

Current chains are source defaults, not proof of the provider that generated every historical output. The CSV’s Gemini version/thinking columns may describe a fallback rather than the primary. The matrix preserves both. The candidate supports text and image input with text output in the endpoint metadata captured for this run. [OpenRouter model](https://openrouter.ai/deepseek/deepseek-v4.1-flash), [reasoning configuration](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens).

| Existing source-default chain | Proposed only after approval |
|---|---|
| gemini-3.5-flash-lite | deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite |
| gemini-3.1-flash-lite | deepseek/deepseek-v4.1-flash → gemini-3.1-flash-lite |
| gemini-3.7-flash → gpt-5.6-terra | deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra |
| gemini-3.7-flash → gpt-5.6-luna | deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-luna |
| gemini-3.7-flash | deepseek/deepseek-v4.1-flash → gemini-3.7-flash |
| gemini-3.1-flash-image | deepseek/deepseek-v4.1-flash → gemini-3.1-flash-image |
| gemini-2.5-flash | deepseek/deepseek-v4.1-flash → gemini-2.5-flash |
| gpt-5.6-luna → gemini-3.7-flash | deepseek/deepseek-v4.1-flash → gpt-5.6-luna → gemini-3.7-flash |

Full use-case/input/output/version/thinking mappings: [replacement matrix](replacement-matrix.xlsx), [service CSV](service-decisions.csv), [selected-row CSV](replacement-coverage.csv).

## Evidence and limitations

The assistant directly read captured inputs and baseline/candidate final outputs and wrote each case judgment. No Python, keyword score, schema score or automated model judge was used for the final quality decisions. JavaScript assembles artifacts and checks transport, file integrity and application compatibility. Per-service review scopes disclose omissions such as unrendered media and website markup.

The requested target remains 5–6 historical staging runs per selected service. Available data provides 159 historical reconstructions across 27 services and 84 fresh pairs across 14 services. Thirty-nine services have six test cases, Product pack briefs has four calls from two plans, and Calendar refresh has five batches from one run. A test case is not necessarily an independent user run. Fresh requests do not satisfy missing historical coverage.

Historical requests were rebuilt using current source prompts and whatever staging inputs were retained. Current profiles, selected assets and inferred request assembly can differ from the original request. Most historical provider/version/thinking settings are unverified; saved output may include user edits. AI director baselines contain aggregate metrics rather than full director output. Native website edit baselines are final documents rather than the original edit plans.

Fresh comparisons use the same constructed prompt and cached input image bytes for both models. These include current public webpage snapshots, representative edit instructions and synthetic over-limit captions. The six caption tests repeat one lab’s saved text; six product refinements reuse one light-edit request and two saved plans. This is useful supplementary evidence, not a historical replay.

Round two adds strict structured-output routing, local contract validation, at most one malformed-response correction, larger completion caps and local image preparation. HEIC inputs are converted with sips to JPEG at quality 92. These adapter/request changes and reconstructed inputs mean this is not a controlled model-only benchmark against round one or staging.

The runner makes up to five DeepSeek attempts. Strict JSON prefers Fireworks, DeepInfra and Morph; eligible errors can switch to JSON mode with DeepSeek/Novita while retaining the same candidate model. Final results used strict mode for 231 cases and text mode for 12; no final JSON-mode case was observed. The existing model chain was not used to conceal a DeepSeek failure.

Earlier Python validators were stopped on the user’s correction and remaining calls were resumed with JavaScript. Completed raw responses were preserved and reviewed directly. Interrupted in-flight calls may have billed without returning usage. Reported cost is therefore a known lower bound for returned round-two attempts, excluding round one, smoke checks and fresh Gemini baselines.

Six native website drafts inflate, prepare, sanitize and validate using actual source functions, with 7–19 recoverable warnings per draft. Three edit operation plans apply and prepare with warnings; three requests ask for clarification. These checks do not establish visual quality, working forms/links, responsive layout or user latency. Images, videos, TTS audio, websites and publishing flows were not generated/rendered end to end.

All 243 result fingerprints match their case objects, all 84 fresh baseline prompt hashes match, and the serialized case-manifest hash matches. All 34 source files in the manifest are unchanged. Every case has one direct written judgment. Complete final answers are inspectable in the local report; private reasoning text and API credentials are excluded.

Scalio Reply, including its Google review replies, is excluded by request. CSV rows marked No/NO, raw-video paths and untested siblings are not approved by this report. Google Business Profile post writing/polishing in scalio-web is included; it is separate from Reply review responses.

## Untested siblings and shortages

- CSV row 9: Initial post variants only. No historical user regeneration request; regeneration behavior is not evaluated.
- CSV row 11: Five batches from one saved refresh run. Need 4–5 more independent runs to meet the requested target.
- CSV row 14: Raw reference-video analysis is outside this text+image test and must retain its current path.
- CSV row 16: Blocked-prompt rewriting is untested. The six fresh comparisons cover narrow storyboard style edits and unsupported shape changes.
- CSV row 19: Pack generation has four product calls from two saved plans; need 3–4 additional independent plan runs. Refinements repeat one softer-light instruction across shots from those plans.
- CSV row 31: Billing-email classification is untested: zero saved provider-billing email events in the inspected staging table. Avatar labels cover only explicitly male prompts.
- CSV row 38: Only the Sora/SD2 enhancement path is covered. Proven-template and other template-adaptation paths remain untested.

The fresh-only services are: Business text analysis, Scraped-photo selection, Brand intelligence, Caption shortening, Copy after photo changes, Trending video ideas, Storyboard style patch, Product pack refinement, Photoshoot ideas, Carousel planning, Click-to-WhatsApp ad copy, Engagement ad copy, Google post polishing, Google post writing. These need historical run capture where the request/output was not retained.

## Proposed next decisions

- Consider only the three scoped candidates for approval: factual attachment descriptions, theme proposals, and narrow storyboard style patches. Candidate means eligible for a scoped discussion, not approved migration or proof of superiority. Theme proposals must not become unverified service claims.
- Collect female/ambiguous gender prompts, unsuitable images and threshold edge cases, accessory classifications, genuinely over-limit multilingual captions, and varied product-refinement instructions before deciding the five limited candidates.
- Keep the other 33 services on their current primary while addressing the specific grounding, duration, locale, image-use and edit-consistency findings. Capture complete input/output/provider/settings/assets for the missing historical runs, then rerun the affected service comparisons.
- Before any approved replacement, implement a service-scoped DeepSeek primary with validated-output and bounded-time fallback. Preserve the existing chain order and each fallback’s own thinking settings. Do not change shared defaults serving excluded features.

## Direct review notes

### Business text analysis

**Hold.** Core categories mostly agree; DeepSeek adds more unsupported details.

Fresh pairs use identical scraped text, but extraction must not turn inference into business facts. WallFusion gains a money-back guarantee and handcrafted production that the supplied text does not establish; logo fallback handling also needs repair.

**Review scope:** Direct reading of all paired extraction outputs and captured website text. No fresh external fact-check or logo image inspection; long BIBA page reviewed for relevant business fields.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Flash-Lite default approximated as disabled; provisional, not a measured compute match.

1. r2-business-text-analysis-78ac063f-9477-4131-ad3a-7591a59f8e9f-fresh — Correct category/logo/products. In-house design is an inference; language english misses prominent Hinglish product names. Gemini detects Hinglish.
2. r2-business-text-analysis-39e7de65-5556-4f66-8acb-18fd89796e75-fresh — Both infer a physical store from an online store. DeepSeek additionally turns satisfaction guarantee into money-back and adds handcrafted production; neither is established in captured text.
3. r2-business-text-analysis-030caa63-c5d4-478a-bffe-1f5d7c2b0f1d-fresh — Correct category, offerings, location and extended hours. Wholesale pricing is a reasonable inference from wholesaler, but should not become a verified pricing promise.
4. r2-business-text-analysis-4222ac1a-dec0-48c8-bdea-20d9ca3b5a20-fresh — Correct clinic classification and exact hours. Both incorrectly accept the explicit generic meta-website-fallback image as a logo; descriptive services are inferred from the clinic name.
5. r2-business-text-analysis-94a24293-b07d-486f-aed0-e3d94fbf9aec-fresh — Broadly plausible BIBA category, collections and brand history. Output tagline is a page/SEO title rather than a demonstrated brand tagline; not independently fact-checked.
6. r2-business-text-analysis-96f38377-f9ed-4a18-ab0e-8b5252455e00-fresh — Grounded jewellery category, shipping and named collections. Services omit saree shapers but cover primary jewellery lines; comparable to Gemini.

### Website visual identity

**Hold.** Useful descriptions, but obstruction guard is incomplete

DeepSeek notices the obstructed pages more clearly than historical outputs. However, the browser-check case still fills brand palette and layout fields using generic interstitial UI despite saying no genuine brand identity is available. Downstream consumers cannot safely rely on the note alone.

**Review scope:** Direct reading of all six inputs and both outputs, plus visual inspection of all six source screenshots.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Source comments: Flash-Lite defaults to zero thinking → reasoning.enabled=false

1. r3-website-identity-c2964559-7c79-43ca-a8f1-3e7eb2de68e5-none-round2 — Faithful navy/blue medical palette and shield description; mixed typography is questionable for the predominantly sans-serif screenshot.
2. r3-website-identity-1ce96b69-62e6-4486-940f-b2a0dd939181-none-round2 — Correctly identifies the blank body and limits evidence to logo/header; good warning missing from baseline.
3. r3-website-identity-c84e560d-e372-428c-a550-fcb6704f6793-none-round2 — Correctly identifies security interstitial, yet returns spinner colors as brand colors and hero-heavy layout. This contradicts the explicit exclusion instruction.
4. r3-website-identity-1e59eb48-0130-4aea-af5c-45f234e7dc77-none-round2 — Captures religious hero imagery and gold/purple palette; says navigation is sans-serif when the visible navigation is serif.
5. r3-website-identity-cba95bb2-a159-47f3-b60c-0e7b433a5f06-none-round2 — Good script-logo/teal/cream/gold reading; typographic detail is mostly useful despite minor wording inconsistency.
6. r3-website-identity-bef25c3b-bae5-4662-a15e-0554be045503-none-round2 — Good dark-teal and serif identity, but calls the visible neon text a cocktail image; historical editorial-grid description fits the split panel better.

### Scraped-photo selection

**Hold.** Comparable selection on product-rich pages; both models accept banners and overstate what metadata proves.

This is a metadata-only task. DeepSeek correctly filters many obvious logos/icons and selects sensible product indices, but its own BIBA reasons explicitly accept navigation/collection banners despite the rejection rule. 9Blings URLs also expose a capture issue: width/height=1 despite large declared dimensions.

**Review scope:** Direct reading of all six candidate metadata pools and both pick lists. Pixels were not part of model input and were not independently inspected for this service.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Flash-Lite default approximated as disabled; provisional, not a measured compute match.

1. r4-scraped-photo-selection-78ac063f-9477-4131-ad3a-7591a59f8e9f-fresh — Sensible spread of tee products plus alternate views. Gemini returns seven distinct hero products; DeepSeek fills ten with duplicates of product families. Both sets are defensible.
2. r4-scraped-photo-selection-39e7de65-5556-4f66-8acb-18fd89796e75-fresh — Avoids the explicit main banner that Gemini selected and picks useful named wall-art products. Scenic ocean and unnamed section assets still need visual verification.
3. r4-scraped-photo-selection-030caa63-c5d4-478a-bffe-1f5d7c2b0f1d-fresh — Excludes the repeated undersized logo and includes ten gallery candidates. Reasons often assert authenticity or interior content without seeing pixels; ranking is weakly evidenced.
4. r4-scraped-photo-selection-4222ac1a-dec0-48c8-bdea-20d9ca3b5a20-fresh — Correctly rejects all avatar assets. Both accept every gallery candidate and cannot establish actual photo quality or relevance from generic labels.
5. r4-scraped-photo-selection-94a24293-b07d-486f-aed0-e3d94fbf9aec-fresh — DeepSeek explicitly selects a navigation banner and collections banner although banners are forbidden. Gemini makes the same mistake; parity is not sufficient.
6. r4-scraped-photo-selection-96f38377-f9ed-4a18-ab0e-8b5252455e00-fresh — Chooses strong product metadata and rejects logos. Selected URLs request one-pixel derivatives despite metadata dimensions; resolve image URL normalization before claiming usable-photo selection.

### Business photo analysis

**Hold.** Vivid concepts, but several factual and classification regressions

DeepSeek produces more specific visual descriptions and usable concepts, but labels a kurta photo as food, invents QR booking, opening status and handcraft/fabric claims. These errors would propagate into downstream marketing.

**Review scope:** Direct reading of all six captured inputs and both outputs, and visual inspection of the six source photos.

**Current chain:** gemini-3.1-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.1-flash-lite  
**Thinking tested:** disabled. Source comments: Flash-Lite defaults to zero thinking → reasoning.enabled=false

1. r5-business-photo-1fb0d6c3-c080-4711-908f-34608db8345d-none-round2 — Good visible description, but converts brand pricing starting at Rs.349 into this exact tee price and invents soft cotton/breathability.
2. r5-business-photo-57bdbd78-0dea-4053-92bc-42ebb9e55c10-none-round2 — Invents first location and opening soon from a storefront image; small-logo OCR is unreliable. Existing output also assumes consultations.
3. r5-business-photo-5c6dd76a-f6c1-4a95-99f4-7540f67b4722-none-round2 — Better respects the apparel profile than the historical café-branded output. However chocolate milk glass is a mistaken visual detail and new outfit assets are only proposed, not present.
4. r5-business-photo-1c9d2483-4ef4-4893-bc8b-f703da20a589-none-round2 — Both infer the dish name; DeepSeek adds unverified hakka noodles, richest curry and bubbling-at-table claims. Product suitability false is questionable for this dominant dish.
5. r5-business-photo-07c8c5bd-2b0f-4cbb-ad0a-9fa5e04b29b1-none-round2 — Calls a clearly visible kurta photo food despite correct garment prose, and claims handmade zari/process evidence without support. Existing output also assumes silk.
6. r5-business-photo-a02c732d-020e-4e6e-ba95-fcef458d8d16-none-round2 — Invents QR appointment booking/check-in and vaccination visit from a reception photo. QR codes alone do not establish their destination or workflow.

### Fashion classification

**Limited candidate.** Broad labels mostly reasonable; confidence logic needs work

Demographic labels match the saved outputs. Western dominance for SMARTLOOK is defensible despite one kurta-like shirt. Captor incorrectly discounts accessories, although this service explicitly includes accessory businesses; do not use its confidence to suppress those businesses without calibration.

**Review scope:** Direct reading of all six inputs and both outputs, plus visual inspection of all 28 source photos.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Source comments: Flash-Lite defaults to zero thinking → reasoning.enabled=false

1. r6-fashion-context-4ff48d8f-2ee4-4c68-9ae0-338c58f61c04-none-round2 — Correctly recognizes fertilizer ads as non-apparel and lowers confidence; 0.35 is still more confident than the baseline 0.1 for irrelevant input.
2. r6-fashion-context-fc4fb3d8-76f3-4ef2-8cde-5d3532882237-none-round2 — Four western-casual examples dominate one kurta-like garment, making western defensible versus historical mixed. The claim that every garment is western is too absolute.
3. r6-fashion-context-596a5ec3-e349-4ae6-87c5-c2f60a002312-none-round2 — Women’s western sandals is a reasonable classification for the five photographed shoe styles.
4. r6-fashion-context-18b4b372-b1ea-40b9-8095-cfb7ddeb1f1d-none-round2 — Western/mixed correctly reflects jackets, romper and a child dress.
5. r6-fashion-context-ebfe6e1e-252a-4d27-ba34-1f112b7e3a7b-none-round2 — Western/mixed correctly classifies the medical scrubs and patient gown range.
6. r6-fashion-context-a6f08e0f-2abb-4b17-8c69-603d5bf31b3b-none-round2 — Western/mixed is usable, but discounts clear accessory evidence and misdescribes the visible red kurta as western. Confidence is not well calibrated for the task.

### Quick photo triage

**Limited candidate.** Comparable on useful-image acceptance

Both models accept six usable references and identify their broad subjects. No rejection examples were sampled, so false acceptance of logos, trackers or irrelevant media remains untested. Reference scores differ enough to affect selection thresholds.

**Review scope:** Direct reading of all six inputs and both outputs, plus visual inspection of all six source images.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Source comments: Flash-Lite defaults to zero thinking → reasoning.enabled=false

1. r6-photo-triage-057057df-970c-4f4b-b466-fedf5c28dd86-none-round2 — Correctly accepts the clear embroidered-kurta lifestyle image.
2. r6-photo-triage-766d79aa-af0b-4611-87e8-0aa25edab22c-none-round2 — Product-on-model is more specific than the saved product label for the child clothing catalogue image.
3. r6-photo-triage-359b2f46-03cd-46d7-9c7a-87f24047a7c9-none-round2 — Correctly identifies the illuminated clock/lamp as the dominant product despite shop clutter.
4. r6-photo-triage-d8b0c9f4-5df1-44c9-a847-752a3b312728-none-round2 — Correctly accepts the clear cake product photo.
5. r6-photo-triage-747415e4-d601-4de8-9ebb-f4b00716fd5e-none-round2 — Correct store-interior label. Lower product/lifestyle scores could suppress useful store imagery; confidence thresholds need calibration.
6. r6-photo-triage-5ad984d5-6192-4820-b341-f676526b9a29-none-round2 — Correct tyre product; high graphic score is understandable with large product branding, but phone screenshot chrome still needs cropping.

### Brand intelligence

**Hold.** More specific personas and voice guidance, but greater unsupported business-claim expansion.

Fresh identical-input comparisons show worthwhile strategy synthesis. Iconic Vision gains single-item ordering, hand finishing, in-house shipping and shorter lead times without evidence. Aainu gains supervised/on-site management. These claims would propagate into downstream content.

**Review scope:** Direct reading of captured inputs and both outputs; generated media is not rendered.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** high. medium → high approximation; no exact Gemini token-budget equivalence.

1. r7-brand-intelligence-715658fa-8a15-486f-aecf-4058ec8438a6-fresh — Good action-led sample sentence and empty trust signals. Accreditation is broadened from laboratories to the centre; preserve the supplied certification scope.
2. r7-brand-intelligence-98fad731-aedb-4a09-becd-9d1c10faa89a-fresh — Strong occasion-led positioning, but adds hand finishing, in-house shipping, single-item acceptance and a middleman lead-time advantage. Gemini remains more restrained.
3. r7-brand-intelligence-7c006ee8-5f0a-4ac6-985e-1c07830f627a-fresh — More grounded bespoke Vastu language than Gemini’s prosperity/energy claims. Proposed personas are clearly strategic inference rather than actual customer evidence.
4. r7-brand-intelligence-73437511-013c-45dc-8492-df4f6d66e9ab-fresh — Correct supplied 4.9/239+ rating and positive themes. Supervised building and on-site staff are not established; location near campus/quiet street is also inferred.
5. r7-brand-intelligence-fc4034f6-e150-4688-948a-0c21efaa8949-fresh — Accurately uses supplied performers/refund/logistics facts. App search-selection-confirmation flow is inferred from the app link, not documented.
6. r7-brand-intelligence-3f45c530-6773-4145-a63e-c7603249366b-fresh — Good view-first voice and supplied services, but sample sentence promises an entirely unchanged view and no harmed bird; these are stronger absolutes than the concrete installation evidence.

### Brand visual system

**Hold.** More concrete and sometimes better grounded in the supplied palette; also more restrictive and less concise.

Several useful art directions, especially Iconic Vision and 10TH COSMOS. However, Max Lab explicitly downgrades a palette labelled real, and Aainu introduces contradictory white-space targets and bans colours present in natural photography. This is not yet a consistently usable short visual brief.

**Review scope:** Direct reading of captured inputs and both outputs; generated media is not rendered.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** high. medium numeric budget 6000 → high; no exact budget equivalent

1. r7-brand-visual-system-715658fa-8a15-486f-aecf-4058ec8438a6-round2 — Overrules the explicit real black/white palette by calling it low priority, adds navy/teal, and makes all text serif. The baseline also adds colours, so it is not a perfect reference.
2. r7-brand-visual-system-98fad731-aedb-4a09-becd-9d1c10faa89a-round2 — Much closer to the stated blue/charcoal palette than the saved pink/rose palette. Engraved-corner motif and in-use memento photography are specific and useful.
3. r7-brand-visual-system-7c006ee8-5f0a-4ac6-985e-1c07830f627a-round2 — Orange/green system follows the supplied logo description better than historical maroon/cobalt. Professional-service scenes are proposed art direction, not evidence of actual clients.
4. r7-brand-visual-system-73437511-013c-45dc-8492-df4f6d66e9ab-round2 — 80% white in palette rules conflicts with 70% in density. No third colour ever enter the frame conflicts with real people, meals and foliage; tighter scope for graphic accents is needed.
5. r7-brand-visual-system-fc4034f6-e150-4688-948a-0c21efaa8949-round2 — Coherent stage photography, saffron accents and hierarchy. More prescriptive than baseline but broadly appropriate; no generated design was checked.
6. r7-brand-visual-system-3f45c530-6773-4145-a63e-c7603249366b-round2 — Installed-net framing and mesh motif are useful. Adds neutral colours and a 12-word caption cap without a need in the supplied brief; a proposed design system, not extracted fact.

### Monthly calendar strategy

**Hold.** Improved structure in several plans; factual invention remains

The plans broadly follow the 22-date window, evergreen emphasis and varied subjects. However, new arrivals, actual production procedures and business-specific advice are asserted without supporting facts. Some historical outputs are also poorly grounded. Festival quality cannot be compared because the historical candidate list was not retained.

**Review scope:** All supplied brand facts, photo summaries, planning constraints, six strategy summaries and all 132 angles from each side read directly. Historical postprocessing/reference allocation omitted from qualitative review.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r8-monthly-strategy-715658fa-8a15-486f-aecf-4058ec8438a6-round2 — Max Lab: useful distinct education and owner-shot planning, but the summary incorrectly says no festivals fall in the window when only the candidate list is empty. The plan stretches accreditation into a guarantee and assumes doctor annotations and cross-location report continuity; the historical plan also adds cold-chain and instant-access claims.
2. r8-monthly-strategy-98fad731-aedb-4a09-becd-9d1c10faa89a-round2 — Iconic vision: useful material, briefing and event-kit variety, but invents newly added designs, zinc-sheet workshop details and every-memento inspection. The old plan also invents newly crafted/in-stock products, coatings and a shipping guarantee. Neither is ready without grounding fixes.
3. r8-monthly-strategy-7c006ee8-5f0a-4ac6-985e-1c07830f627a-round2 — 10TH COSMOS: less expansive than the saved industrial/luxury service claims, yet invents a site-survey-to-written-strategy workflow and real most-asked questions. Several plot selection/alignment topics overlap despite claims of uniqueness.
4. r8-monthly-strategy-73437511-013c-45dc-8492-df4f6d66e9ab-round2 — Aainu PG: substantially more restrained than the saved round-the-clock security, single-room and review claims absent from this reconstruction. Useful moving and housemate advice; quiet/safety/meal positioning mostly follows supplied photo summaries, which themselves need factual confirmation. Broadly usable concepts with that limitation.
5. r8-monthly-strategy-fc4034f6-e150-4688-948a-0c21efaa8949-round2 — Shobhnam: covers the supplied services and refund terms with reasonable planning advice. “Booking without the risk” overstates a cancellation refund; coordinator workflow and app browsing are inferred. Some checklist/briefing topics are close. No made-up discount or phone number.
6. r8-monthly-strategy-3f45c530-6773-4145-a63e-c7603249366b-round2 — Bird Net: covers the real seven-service range and improves on saved UV-resistance and spike-material inventions. Still presents inferred fitting workflow, personal experience lessons and site coordination as actual business practice; some safety/view promises need more careful qualification.

### Calendar variants

**Hold.** Better null handling and palette compliance; unsupported claims and narrow coverage

All six requests are for Max Lab and reconstruct initial generation without prior variants or regeneration instructions. DeepSeek repeatedly invents a link-in-bio booking route, strengthens accreditation into blanket trust claims and asserts local fever trends without evidence. The outputs do not establish safe regeneration behavior.

**Review scope:** Full captured brief and slot differences, all six output content fields and slide briefs read directly; large downstream assembled image prompts were excluded. No medical claims were independently validated or graphics rendered.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r9-calendar-variants-4c57a517-c18c-4d23-bfaa-607ba6152901-round2 — Single educational: a useful three-point checklist and correct black/white palette, but “single biggest signal” and booking via link in bio are unsupported. The older report-access post also overstates instant access in its explanation.
2. r9-calendar-variants-491fb57d-37dc-498a-a7ba-94b7000f470e-round2 — Single promotion: follows the home-collection USP and removes irrelevant video fields, but changes accredited laboratories into this centre being an accredited laboratory and invents link-in-bio booking. Some collection coverage is inferred.
3. r9-calendar-variants-08aadaf7-b853-4f79-b746-2fd80388fd3b-round2 — Carousel: plausible three-step patient journey and four slides. The final slide emphasizes accreditation instead of clearly showing the third report-download step. Online/link-in-bio booking is inferred; change from lipid education is allowed because no original angle was supplied.
4. r9-calendar-variants-1d1722e3-f5f2-4bf0-ab28-781fafefba04-round2 — Seasonal: asserts late September is peak local fever season and that cases are not easing, with no epidemiological evidence. “A blood test tells you the difference” is overconfident. The old liver/kidney seasonal-recovery framing is also unverified.
5. r9-calendar-variants-b9d7a944-1b5e-49a6-b12f-99837dd60fd3-round2 — Carousel: readable home-collection story, but invented booking-in-minutes and fast reports. It nearly repeats the other carousel because neighboring variants were absent from the reconstruction; do not attribute missing-context diversity failure solely to the model.
6. r9-calendar-variants-908d6745-0c32-4f22-85eb-bae2ef38bd95-round2 — Single educational: clear accreditation explainer, but invents a marked/circled NABL report image and describes accreditation as the difference between trustworthy and untrustworthy numbers. Booking link remains unsupplied.

### Caption shortening

**Limited candidate.** Comparable; cleaner newline handling in one pair

All six retain the main message, CTA and hashtags while removing repeated blocks. This is a weak compression challenge made from one lab’s repeated captions, not six real shortening requests; diverse nonrepetitive and multilingual captions remain untested.

**Review scope:** Direct reading of all six complete supplied captions and both outputs. Judgment is preservation/compression, not verification of the original medical claims.

**Current chain:** gemini-3.7-flash → gpt-5.6-luna  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-luna  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r10-caption-shortener-4c57a517-c18c-4d23-bfaa-607ba6152901-fresh — Preserves the online-report hook, location, CTA and hashtags; minor wording compression is faithful.
2. r10-caption-shortener-491fb57d-37dc-498a-a7ba-94b7000f470e-fresh — Preserves all useful content and emits real line breaks where Gemini emitted literal backslash-n sequences.
3. r10-caption-shortener-08aadaf7-b853-4f79-b746-2fd80388fd3b-fresh — Keeps the three explanatory items and CTA while trimming wording; no new medical claim is added by the shortening.
4. r10-caption-shortener-1d1722e3-f5f2-4bf0-ab28-781fafefba04-fresh — Faithfully deduplicates seasonal copy, though removing the article before heavy monsoon slightly worsens grammar.
5. r10-caption-shortener-b9d7a944-1b5e-49a6-b12f-99837dd60fd3-fresh — Keeps the intended pre-test message; compression makes the chai/gym sentence less elegant than Gemini, but preserves meaning.
6. r10-caption-shortener-908d6745-0c32-4f22-85eb-bae2ef38bd95-fresh — Correctly removes repeated hooks and blocks while retaining the medical-guidance qualification and CTA.

### Copy after photo changes

**Hold.** Mixed; often preserves more of the original than Gemini, but one adaptation makes the copy worse.

Six fresh tests all replace references with the same Max Lab logo. This is narrow coverage of an image-change task. Most copies remain on theme; the monsoon version becomes repetitive and turns generic pathology services into a seasonal recovery check-up. The baseline is clearer in that case.

**Review scope:** Direct reading of captured inputs and both outputs; generated media is not rendered.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r10-refs-copy-adaptation-4c57a517-c18c-4d23-bfaa-607ba6152901-fresh — Keeps the original copy unchanged; this is defensible for a logo-only reference and preserves voice better than an unnecessary rewrite. Existing claims are inherited, not newly verified.
2. r10-refs-copy-adaptation-491fb57d-37dc-498a-a7ba-94b7000f470e-fresh — Removes the original instant-online-access promise and preserves the checkup/home-collection theme; a useful conservative edit.
3. r10-refs-copy-adaptation-08aadaf7-b853-4f79-b746-2fd80388fd3b-fresh — Preserves the lipid-guide copy and CTA verbatim. Appropriate for logo-only change; not an independent verification of medical explanations.
4. r10-refs-copy-adaptation-1d1722e3-f5f2-4bf0-ab28-781fafefba04-fresh — Adds an awkward repeated NABL phrase and implies a named seasonal recovery check-up that the allowed services do not establish. Gemini is shorter and more usable.
5. r10-refs-copy-adaptation-b9d7a944-1b5e-49a6-b12f-99837dd60fd3-fresh — Minimal brand-name expansion preserves the locked educational topic. Clinical-precision wording is inherited; neither output validates that claim.
6. r10-refs-copy-adaptation-908d6745-0c32-4f22-85eb-bae2ef38bd95-fresh — Almost unchanged, with relevant brand hashtags. Preserves the topic and CTA but this does not test a materially different reference photo.

### Calendar refresh

**Hold.** Better literal photo adherence, but mixed inputs remain incoherent and new claims appear

Five batches are one independent staging run, not five runs. The source mixes a dot business name, fashion category, software-security claims and unrelated winery/preschool/fashion references. DeepSeek uses the assigned descriptions more literally and often preserves copy, but produces incoherent combinations and newly invents a seasonal phishing claim. Missing previous facts also prevent reliable removed-service evaluation.

**Review scope:** Current facts, all 20 source posts and photo descriptions, and both sets of content changes read directly. Repeated identical fields were displayed once; downstream assembled image prompts were excluded. One independent source run only.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** high. medium → high approximation; no exact compute equivalence.

1. r11-calendar-refresh-fb548638-6557-40e3-aa31-cc0d4913b08d-a69e7160b3a6c0a80a36e93747f87d9d6cbbf8fbb2f7cbcc9dffd5aa731dc567-round2 — Batch 1: preserves much source copy and appends CTAs, but uses winery branding for a hosting launch, jewellery for abuse prevention and a lake/preschool alternative for API security. Instant protection and hosting claims are inherited; newly rewritten peak-load text is unnecessary.
2. r11-calendar-refresh-fb548638-6557-40e3-aa31-cc0d4913b08d-c3043bced7c959befee39aa9a9c5a9e7a343e6e5c23d454ac4efedc2a8a972a3-round2 — Batch 2: introduces “monsoon ... rise in phishing and malware attempts” without evidence while replacing uptime messaging. Other posts preserve source text but offer unrelated wedding/wine/children images even when a developer laptop is available.
3. r11-calendar-refresh-fb548638-6557-40e3-aa31-cc0d4913b08d-43b025a40ee0432dde206dbef3793fc1b1b787a00c7f5f18bf490707c7fdd43d-round2 — Batch 3: removes some AI/hosting references although previousFacts is null and the prompt says not to infer removal from incomplete data. Source-inherited absolute savings/security claims remain. A hair-oil bottle is offered as a product-focused alternative for infrastructure messaging.
4. r11-calendar-refresh-fb548638-6557-40e3-aa31-cc0d4913b08d-6cd95073fa3b7cdcdbae871bf1a36c3b4315e4870cea20764e20288e069907c4-round2 — Batch 4: keeps carousel sequencing and original English hashtags, unlike corrupted mixed-script hashtags in several saved outputs. However, every slide forcibly pairs technical material with wine, preschool or event imagery. Some irrelevant video/shot fields remain; this is not a usable marketing result.
5. r11-calendar-refresh-fb548638-6557-40e3-aa31-cc0d4913b08d-ad081f8dea6571bfeff868d1f9b50603cca3f59a0120c0c62b23b057adee9368-round2 — Batch 5: correctly includes the explicit Instagram CTA and clears irrelevant media fields. It still pairs moderation claims with beach fashion, toddlers and a banquet, and preserves the inherited unsupported promise that automation stops threats instantly.

### Batch video ideas

**Hold.** Useful variety, with unsupported claims in both candidate and saved batches; locale and photo-context reconstruction prevent a controlled winner claim.

The six complete 20-idea batches are readable and often actionable, but DeepSeek invents timings, founder history and customer outcomes, repeats angles, and sometimes requests an on-camera host while presenter is none. Source fact contamination and incomplete historical photo ordering also need correction before migration.

**Review scope:** Direct reading of all six reconstructed business/photo-description inputs, common source system prompt and all 120 saved plus 120 DeepSeek ideas. Original full candidate-photo order is not preserved; current profiles and selected photo descriptions substitute for historical context. No videos rendered and no claim of image-by-image reinspection in this final review.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r12-video-ideas-batch-bd171e31-a23a-4b6e-9e6e-671116f7c720-round2 — Graphiora: good use of the cat and panther drawings, but “one whisker takes ten minutes,” founder motivations and commission turnaround/feedback promises are not supplied facts. Pet-portrait commissions are inferred, and the Prabhakar_atelier flyer conflicts with Graphiora. Both models extrapolate processes; repeated cat/panther angles limit diversity.
2. r12-video-ideas-batch-8e396864-1ea3-4d96-9127-676082f12dd4-round2 — Dev Babu Jewellery: useful styling and neckline/tikka angles, but many ideas repeat the complete coordinated-set pitch. “Every order contains” generalizes one pictured set to all orders; Full Set vs Mixing calls for a presenter despite presenter:none. Existing ideas also repeat and infer materials. Five available photos cannot meet a literal twenty-distinct-photo requirement.
3. r12-video-ideas-batch-d57ef21c-0fa9-42ba-b209-8bed24c46c3a-round2 — Fortcitytaxi: the supplied 24/7 service, fleet and transparent rates are used well. However “first booking takes one minute,” “one call” booking, “we are not an app,” and years of driver tenure are unsupported. A rider quotation is described as dramatized but is placed under customer proof; any output would need a visible dramatization label. Historical hooks largely use a different language from the reconstructed English profile.
4. r12-video-ideas-batch-675a3844-9a47-4c0a-b8d6-ee7c2020662c-round2 — 13Trips: the 97% aggregate track record is supplied, but “he was rejected once” fabricates the pictured France client’s prior refusal. “Small documentation mistakes cause most visa refusals” and a more-than-twice-a-year membership threshold have no evidence. Some goals combine enquiries and follower growth despite the one-goal instruction. The existing batch is also unsafe: it invents yesterday, zero application errors and an individual 97% success probability. Neither should be treated as verified visa guidance.
5. r12-video-ideas-batch-8d42a3cd-18bf-4755-ad5c-a3dbcd9f1965-round2 — Svolto: practical clock/timer/weather concepts and clearer variety, but “fully set up in about a minute,” recent improvement walkthroughs and product longevity beyond the supplied commitment are unverified. A maker/founder story is invented; “no glasses needed” overpromises legibility. Existing output also invents zero eye strain, zero glare and product details. Two photos make unique image assignment impossible, so reuse alone is not a model failure.
6. r12-video-ideas-batch-52264bae-54d0-4222-9916-f666ee695074-round2 — Royal Collection: DeepSeek sensibly prioritizes real clothing photos over the contaminated social-platform USPs and avoids the existing baseline’s exact-size/stock guarantees. It still invents a founder motivation, current evening opening/availability and new-arrival framing, and two ideas have presenter:none while directing a host. The old batch invents daily opening and recurring fresh stock. Current inventory and store history require confirmation.

### Onboarding video idea

**Hold.** DeepSeek follows the reconstructed photo descriptions more closely; historical outputs often refer to different products, so these are not reliable model wins.

The supplied hero descriptions and current profiles sometimes differ materially from saved outputs. DeepSeek produces focused 20-second concepts but can add unsupported product properties or review content. Resolve replay fidelity and these claims before replacing.

**Review scope:** Read all six reconstructed profiles, photo descriptions and outputs. No image pixels were supplied to this text call. Historical hero/profile mismatches prevent a fair winner count.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r12-onboarding-video-idea-70d1d89e-fee3-4fe8-ac06-08f6a02c2b67-round2 — Local store concept is grounded in the supplied Surat address. It avoids the saved output’s unsubstantiated authentic/trusted-partner claims.
2. r12-onboarding-video-idea-ffce01a9-a554-4ce6-aa96-421eaa981137-round2 — Magenta/gold description matches supplied text better than historical red/silver; silk is inherited from the description rather than visually established. Profile-link CTA is not supplied.
3. r12-onboarding-video-idea-695cbb1c-660b-4aba-a7db-78b356236294-round2 — Preserves English and the actual app utility. Proposes showing real store reviews without supplied review content; needs a verified review asset.
4. r12-onboarding-video-idea-0512f7e2-4c1a-4875-a2ae-54e5aca9dc91-round2 — Uses the supplied bracelet rather than unrelated historical butterfly earrings. Affordability/topaz claims depend on generated photo metadata, with no business profile to verify them.
5. r12-onboarding-video-idea-632d83af-875b-431b-bfe2-9d7db49891c9-round2 — Uses the supplied crane/lotus CNC panel rather than historical peacock mirror art. Custom-order positioning is inferred, so historical comparison is inconclusive.
6. r12-onboarding-video-idea-75986d43-bfd4-4d36-b0c9-0a407fd8d3b8-round2 — Lip-balm concept is more relevant than the saved skincare-consultation ad. A tinted swatch and dessert-sweet product property are not established by the supplied description.

### Trending video ideas

**Hold.** More distinctive creative angles than Gemini, with materially worse factual grounding in several cases.

Different formats are recognizable and varied, but the restaurant gets fictitious regular-customer behavior and a hundred hot-delivered boxes, the shop gets invented opening times/process durations, and the clinic gets an invented four-second vaccination and dawn opening. Creativity does not make these publishable.

**Review scope:** Direct reading of all six format briefs, business facts, photo descriptions and all 18 ideas from each model. No generated videos; source prompt/schema duration conflict is reported separately.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** high. high, matched to source primary high.

1. r12-trending-video-ideas-1fb0d6c3-c080-4711-908f-34608db8345d-fresh — Three distinct, culturally specific brand films. Some descriptions exceed the requested brevity; first goal bundles several actions and the second hook mixes language. Price is correctly qualified as starting price.
2. r12-trending-video-ideas-57bdbd78-0dea-4053-92bc-42ebb9e55c10-fresh — Good BTS mechanics but invents 9 a.m., three-hour shelf arrangement and checking every product. Existing baseline also infers routine care; DeepSeek adds explicit unsupported specifics.
3. r12-trending-video-ideas-5c6dd76a-f6c1-4a95-99f4-7540f67b4722-fresh — Both models invent apparel/products/packaging while attaching café-photo indices. DeepSeek adds a production-process claim; reject the conflicting inputs instead.
4. r12-trending-video-ideas-1c9d2483-4ef4-4893-bc8b-f703da20a589-fresh — Contrarian hook disparages the restaurant’s noodles and invents what regulars buy. A hundred boxes, every one hot, and parents asking a quoted question are fabricated customer/process evidence.
5. r12-trending-video-ideas-07c8c5bd-2b0f-4cbb-ad0a-9fa5e04b29b1-fresh — Recognizable UGC concepts, but hand-detailed embroidery and scripted customer/mother approval are presented as real proof. Need clear dramatization and owner-approved claims.
6. r12-trending-video-ideas-a02c732d-020e-4e6e-ba95-fcef458d8d16-fresh — Good abstract identity direction, but four-second procedure and dawn opening are invented. The first concept also assumes the interior is literally pink/blue despite a purple-bench reference.

### Storyboard compilation

**Hold.** Plans are structured, but grounding, asset references and pacing still fail

DeepSeek produces coherent beat structures, but adds real-world proof claims, omits required asset references and exceeds spoken word budgets in several beats. The source prompt itself encourages unverified proof/caveats and has conflicting asset-description rules, so fixing only the provider will not solve this.

**Review scope:** Full six briefs, source instructions, observed asset metadata and complete beat/style text from both sides read directly. Video generation, speech timing in audio, identity continuity and final renders are untested.

**Current chain:** gemini-3.7-flash → gpt-5.6-luna  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-luna  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r14-storyboard-70d1d89e-fee3-4fe8-ac06-08f6a02c2b67-round2 — AUSPORA: invents every product being sourced directly from trusted brands. The interior shelf beat has no asset reference, rotates a labelled bottle and exceeds the eight-second dialogue cap. The old plan also invents serums and an interior; neither is reference-faithful.
2. r14-storyboard-23a306f0-80c1-4bc3-ab9e-55d32a8d468c-round2 — Cat portrait: preserves a recognizable drawing-process arc and hand-only presence, but the eight-second outline/fur line is too long. It invents finger-smudging as the real technique and puts caption-free emphasis on the macro rather than the final reveal. Finished-art fidelity still needs rendering.
3. r14-storyboard-ffce01a9-a554-4ce6-aa96-421eaa981137-round2 — Promotion: explicitly shows the reference model while declaring presenter forbidden and product-only, conflicting with the no-identifiable-people contract. It keeps asset references but repeats visible garment descriptions and the silk assumption from conflicting metadata. The baseline also adds pure silk and a new collection.
4. r14-storyboard-d2996fce-e5dc-4ec5-b305-be72d885c005-round2 — Dev Babu necklace: stronger on-camera/b-roll shape, but fabricates daily visitors, weekend rush and a main-bazaar landmark. The final presenter beat omits the product reference although the necklace should remain worn. Store details are invented. The old output also uses an unsupported landmark.
5. r14-storyboard-d8af8af3-30d7-4893-a37d-bb5b32b75e98-round2 — Panther art: follows the print-versus-handmade concept, but closes the sketchbook and hides the hero in the final beat. The eight-second macro narration exceeds its cap. Both sides elaborate graphite texture from the supplied description; no rendered fidelity evidence.
6. r14-storyboard-fd9eebf3-cbc2-4a68-bce4-ec5d123e6ebd-round2 — Commission flyer: presents the portfolio clearly and avoids the old added organic claim, but the eight-second line is over budget. Both sides advertise Graphiora while the supplied flyer is branded prabhakar_atelier; this unresolved source conflict limits a fair replacement decision.

### Storyboard style patch

**Candidate.** Comparable

Both models make the requested narrow style edits and refuse the two changes that would alter video shape. DeepSeek gives useful compiled descriptions; no broad quality advantage is established.

**Review scope:** Both complete outputs read directly; actual local application also confirms scene timing/order/text and unrelated style fields are preserved. Rendered appearance is not tested.

**Current chain:** gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r16-storyboard-patch-70d1d89e-fee3-4fe8-ac06-08f6a02c2b67-fresh — Plain studio request: DeepSeek supplies a neutral studio, simple cosmetic props and matching compiled description. Gemini is similarly usable.
2. r16-storyboard-patch-23a306f0-80c1-4bc3-ab9e-55d32a8d468c-fresh — Soft-daylight request: both revise the visual grade and compiled paragraph while keeping the art-desk concept.
3. r16-storyboard-patch-ffce01a9-a554-4ce6-aa96-421eaa981137-fresh — Add-scene request: both correctly return an empty changed list.
4. r16-storyboard-patch-d2996fce-e5dc-4ec5-b305-be72d885c005-fresh — Remove-scene request: both correctly return an empty changed list.
5. r16-storyboard-patch-d8af8af3-30d7-4893-a37d-bb5b32b75e98-fresh — Tidy-showroom request: both provide a credible showroom description. DeepSeek is more detailed, but that is a style difference rather than a clear win.
6. r16-storyboard-patch-fd9eebf3-cbc2-4a68-bce4-ec5d123e6ebd-fresh — Warm-evening request: both preserve the curator/gallery setup and apply evening lighting in the compiled paragraph.

### Image prompt enhancement

**Hold.** Clear intent regression on sitting-pose request

The model can produce useful detailed prompts, but it converts a sitting-pose request into trousers on a hanger or flat surface with no person. Other cases introduce material guesses and copy claims. For an enhancer, preserving the requested transformation is essential.

**Review scope:** Direct reading of all six text inputs and both outputs. This service’s source-image sets were not fully visually adjudicated; the sitting-pose failure is evident from the request/output alone.

**Current chain:** gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r18-image-enhancement-b33e3b27-666c-4495-a10d-5e0a0ee8cc85-round2 — Both outputs offer event packages. The replay declares a final official logo but contains only two images despite a two-image product/reference contract; logo provenance needs correction before this branch is approved.
2. r18-image-enhancement-57918b60-164d-4df9-b6af-96a121c5767c-round2 — DeepSeek narrows the bracelet to pink rose-quartz-like beads while the baseline mentions pink and green. Needs source-photo adjudication for whether both references should be preserved.
3. r18-image-enhancement-24c22802-d129-490e-a28c-68bc4df4076e-round2 — Preserves seated-model intent and strong garment constraints, but adds guessed cotton/cotton-linen composition that was not in the text brief.
4. r18-image-enhancement-5f46211c-3721-4c95-b80f-a41906dd1083-round2 — Fails the sitting-pose intent by explicitly requesting no people/no mannequin and suggesting a hanger or flat surface. Baseline’s seated waist-down model is more useful.
5. r18-image-enhancement-a5f42ba4-70b0-4ae7-9486-2e55107c22d6-round2 — Front-only bottoms is ambiguous; DeepSeek chooses a hanger whereas baseline chooses a wearer. Cannot infer owner preference from this terse prompt.
6. r18-image-enhancement-6059fd0b-325b-4ff5-9e70-81e9f529e795-round2 — Front-view product treatment is plausible, but horizontal panel seams and material are visual claims needing source-photo verification before approval.

### Product pack briefs

**Hold.** Some better grounded Amazon copy, but uneven shot precision and insufficient independent examples

Four outputs cover only two product plans. DeepSeek removes some unsupported material and dimension claims in the saved barbell plan, but the Flipkart prompts lose concrete camera distances and occasionally request unseen wear. New angles, scaling and product identity need generated-image evaluation before adoption.

**Review scope:** Complete plan inputs and both textual shot sets read directly; original product photos were inspected earlier. No new product images were generated or judged.

**Current chain:** gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r19-product-pack-93104064-970a-4440-a20b-0e47e14eaccf-8ab01e2e-38bc-405d-a0fc-7a7647bfc228-round2 — Amazon barbell: seven requested shots are supplied; the extra box shot reflects the current prompt and is not a model win over the older six-shot baseline. Visible 30LB/Synergee labels are better grounded than historical steel-core/rubber and dimension claims. Scale remains approximate without measurements.
2. r19-product-pack-93104064-970a-4440-a20b-0e47e14eaccf-ab09d96d-61d6-4288-a2d8-01d7b57e182e-round2 — Amazon motorcycle: varied settings and views preserve the main white-background shot, but some distances are unspecified, the dark alternate setup risks poor separation, and sun-behind-camera with shadows toward the foreground is contradictory. New-side mechanical details need image review.
3. r19-product-pack-9ce583b2-56c1-4389-b9ab-0b61e415d4e2-8ab01e2e-38bc-405d-a0fc-7a7647bfc228-round2 — Flipkart barbell: preserves the product and no-people constraints, but generally omits the concrete camera distances required by the prompt. The extreme macro is vague about which useful construction detail it should reveal. Props establish only approximate scale.
4. r19-product-pack-9ce583b2-56c1-4389-b9ab-0b61e415d4e2-ab09d96d-61d6-4288-a2d8-01d7b57e182e-round2 — Flipkart motorcycle: the Indian-lane context fits the brief, but the craft shot invents micro-scratches and uses a dark setting despite the brighter marketplace direction. Most shots omit precise distance; unseen-side identity is untested.

### Product pack refinement

**Limited candidate.** Comparable; DeepSeek preserves more of the original camera and background instructions.

All six fresh tests apply the same softer-light edit to shots from only two saved plans. The requested lighting change is handled without moving the camera or changing required white backgrounds. Useful evidence for this narrow edit, not for arbitrary refinements.

**Review scope:** Direct reading of all six supplied prompts and both outputs. Product identities had been visually checked in product analysis; no new images were generated. Two saved plans, one repeated edit request.

**Current chain:** gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r19-product-pack-refinement-93104064-970a-4440-a20b-0e47e14eaccf-8ab01e2e-38bc-405d-a0fc-7a7647bfc228-main-fresh — Preserves 30-degree angle, three-foot distance, 85% fill and pure white; softens contact shadows. More complete constraint retention than Gemini.
2. r19-product-pack-refinement-93104064-970a-4440-a20b-0e47e14eaccf-ab09d96d-61d6-4288-a2d8-01d7b57e182e-main-fresh — Preserves the four-metre motorcycle side view and white background. Soft edge definition/minimal contrast may reduce separation and should be checked in a rendered image.
3. r19-product-pack-refinement-9ce583b2-56c1-4389-b9ab-0b61e415d4e2-8ab01e2e-38bc-405d-a0fc-7a7647bfc228-main-fresh — Preserves elevated angle, distance, fill and the policy-background sentence while replacing crisp light with feathered shadows.
4. r19-product-pack-refinement-9ce583b2-56c1-4389-b9ab-0b61e415d4e2-ab09d96d-61d6-4288-a2d8-01d7b57e182e-main-fresh — Keeps the full side profile, three-metre distance, white backdrop and exclusion of props/reflections.
5. r19-product-pack-refinement-93104064-970a-4440-a20b-0e47e14eaccf-8ab01e2e-38bc-405d-a0fc-7a7647bfc228-alt-fresh — Keeps the straight-on full-length view, concrete and slate background. Neither original nor DeepSeek states numeric frame fill here; Gemini adds 30%.
6. r19-product-pack-refinement-93104064-970a-4440-a20b-0e47e14eaccf-ab09d96d-61d6-4288-a2d8-01d7b57e182e-alt-fresh — Keeps the low rear three-quarter camera and dark setting. Softens the lighting without changing the scene; camera summary should describe the camera rather than lighting.

### Photoshoot ideas

**Hold.** Mostly practical fashion concepts, but weak category handling for storefront and clinic photos.

Fresh tests show the model can create coherent outfit settings and preserve common garment details. It invents an AUSPORA fashion collection and turns a clinic into an ethnicwear set, with an unverified Ayushman Bharat story. The clinic Gemini retry stays closer to a healthcare theme, but also invents WHO posters, staff identities and an event. Both require better grounding; the original 503 is archived separately.

**Review scope:** All six input photos, source directions and all 30 ideas from each model were read directly. The clinic Gemini baseline succeeded on retry after an initial 503. No generated photos were evaluated.

**Current chain:** gemini-3.1-flash-image  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.1-flash-image  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r20-photoshoot-ideas-1fb0d6c3-c080-4711-908f-34608db8345d-fresh — Keeps the cream pants more consistently than Gemini’s skirt/jeans substitutions. The mountain idea’s rear/side styling hides a front-printed graphic; first three directions are usable.
2. r20-photoshoot-ideas-57bdbd78-0dea-4053-92bc-42ebb9e55c10-fresh — Keeps the actual storefront but invents the store’s fashion/ethnicwear merchandise. Repeating transcribed tax-ID text introduces unnecessary OCR risk; baseline also invents stock/interior scenes.
3. r20-photoshoot-ideas-5c6dd76a-f6c1-4a95-99f4-7540f67b4722-fresh — Coherent café concepts anchored to the actual table items; strong preservation language. More people/fashion-led than needed, but proposed styling is within the creative task.
4. r20-photoshoot-ideas-1c9d2483-4ef4-4893-bc8b-f703da20a589-fresh — Five varied settings with the curry central. Spice/ingredient-origin narrative is not established by the image, and adding cream changes the exact garnish; baseline also invents ingredient/process stories.
5. r20-photoshoot-ideas-07c8c5bd-2b0f-4cbb-ad0a-9fa5e04b29b1-fresh — Coherent worn-garment catalogue/heritage/lifestyle ideas, with useful front-view option. Silk and precise embroidery metal should be left to the reference, not asserted from appearance.
6. r20-photoshoot-ideas-a02c732d-020e-4e6e-ba95-fcef458d8d16-fresh — Clinic: Gemini retry succeeded after the preserved original 503. DeepSeek turns the clinic into ethnic-fashion catalogue sets and invents an Ayushman Bharat poster and health-card campaign. Gemini keeps a clinic/community theme, but also invents WHO posters, staff identities and an event; neither is factual ground truth.

### Attachment descriptions

**Candidate.** DeepSeek generally more useful

The descriptions capture actual poster text, layout and screenshot context more usefully than the saved broad scene labels. Small descriptive inaccuracies remain, but no material invented business offering appeared in this sample.

**Review scope:** Direct reading of all six inputs and both outputs, plus visual inspection of all six source images.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Source comments: Flash-Lite defaults to zero thinking → reasoning.enabled=false

1. r22-attachment-description-661e19e1-2130-48d7-b098-2eea45269e51-none-round2 — Correctly reads COMING SOON and WIRELESS HEADSET; accurate silver/black headphones and night street scene.
2. r22-attachment-description-fb39e61a-a427-472b-9961-a402935af8b3-none-round2 — Correctly identifies the branded school poster and its real headline/subject list; reference is a better creative role than generic scene.
3. r22-attachment-description-e5bbb80b-604a-4bf9-89bd-3763ac3039a4-none-round2 — Correct festival and Flyzone/Ahmedabad context from visible text; suitable reference description.
4. r22-attachment-description-7b661600-30e5-4318-9b50-5708a0bcdaf9-none-round2 — Faithful description of the pregnancy infographic without adding specific medical instructions.
5. r22-attachment-description-d5d61bac-f5f6-4eeb-abae-28f545f956fe-none-round2 — Correctly includes paused-video screenshot and play overlay that the baseline omitted; useful for avoiding reuse of UI chrome.
6. r22-attachment-description-5d8a7ed0-07d1-4d6a-9975-77d43b0e7e17-none-round2 — Correct Ruby Green’s Plus Five product identification. The exact body-point wording is slightly over-specific, but the product and packaging are recognizable.

### Carousel planning

**Hold.** Coherent four-slide structure, but DeepSeek is more prone to inventing proof and products.

Fresh image pairs expose unsupported cotton/breathability/screen-printing claims, invented KPD garments, a fusion-menu story and WHO references absent from the source. Several existing outputs also overclaim; neither should be treated as ground truth.

**Review scope:** Direct reading of all six goals and paired slide plans. All six source images had been visually inspected; no generated carousel images were rendered.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Flash-Lite default approximated as disabled; provisional, not a measured compute match.

1. r22-carousel-planning-1fb0d6c3-c080-4711-908f-34608db8345d-fresh — Good visual match in the opening. Soft cotton, all-day breathability and screen-printing process are unverified; multiple other graphic tees are invented from one photo.
2. r22-carousel-planning-57bdbd78-0dea-4053-92bc-42ebb9e55c10-fresh — More restrained than the baseline’s dermatologist recommendation, but popular/trusted labels and a stocked interior are still inferred from a storefront.
3. r22-carousel-planning-5c6dd76a-f6c1-4a95-99f4-7540f67b4722-fresh — Reconciles an apparel profile with a café photo by inventing a clothing range, cotton weight and lasting construction. This hides the source-data conflict.
4. r22-carousel-planning-1c9d2483-4ef4-4893-bc8b-f703da20a589-fresh — Adds a fusion positioning, noodles, spice ingredients and fresh-food claims from a curry photo and Chinese Restaurant category. Neither source proves the actual menu story.
5. r22-carousel-planning-07c8c5bd-2b0f-4cbb-ad0a-9fa5e04b29b1-fresh — New look and silk/comfort claims are not established. Final image prompt offers incompatible alternatives instead of a concrete composition.
6. r22-carousel-planning-a02c732d-020e-4e6e-ba95-fcef458d8d16-fresh — WHO guidelines are not visible in the source poster; invents a clinician portrait and consultation scene without an approved identity. Baseline stays closer to the facility.

### Native website generation

**Hold.** Copy is often more restrained than staging, but factual and visual gaps remain

All six drafts prepare and validate after application recovery, with 7–19 warnings per draft. Copy is generally more concrete and less inflated than the saved sites, but unsupported product properties, stock/response promises and unsuitable media remain. Historical asset subsets differ; layout, accessibility and live interactions were not visually evaluated. Long retry-inclusive completion times also need attention.

**Review scope:** Six business seeds, core asset descriptions, all baseline copy values, all DeepSeek copy entries/SEO and media selections read directly. Nested AI marketing suggestions and full HTML/CSS were not exhaustively judged. Actual source preparation/validation ran on every draft; no rendered website comparison.

**Current chain:** gemini-3.7-flash → gpt-5.6-luna  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-luna  
**Thinking tested:** high. medium → high approximation; no exact compute equivalence.

1. r23-native-website-generation-f532b598-e5d8-40e5-995a-12902c77bad7-round2 — Himalayan Restaurant: substantially less invented menu detail than the baseline. Still promises a waiting seat even for walk-ins and selects a supplied stock portrait explicitly described as unrelated to the restaurant. Live music and buffet are supplied facts. Contact/privacy claims need owner verification.
2. r23-native-website-generation-d12316c3-376b-44b1-93a8-c5a5f83aa968-round2 — Shining Baby World: uses actual photographed model names and suggests checking stock, improving on the baseline’s invented pickup/assembly guarantees. Yet adds training wheels and multiseason durability; “on the floor now” conflicts with its own historical-photo caveat. Both outputs are mostly English despite Hinglish locale.
3. r23-native-website-generation-e3350be0-060a-4a40-8422-81f6a27c7cbc-round2 — Sri Kirthi: aligns the catalogue to actual saree and kurta descriptions and preserves supplied hours. Invents colour retention and a relationship between the K logo and saree border. The saved site has much larger invented departments and literal privacy placeholders; that does not make the new draft ready.
4. r23-native-website-generation-669570c5-2376-4795-90da-f26c3ad550f2-round2 — Korean Kick: grounded main food categories and supplied opening hours, with less invented menu variety than the baseline. Still makes made-to-order and fast-reply/table-readiness promises beyond the verified profile, and writes mainly English for Hinglish locale. Food names/preparation require owner confirmation.
5. r23-native-website-generation-dc8e190b-bb24-4081-ba1f-7430d50902cc-round2 — Chopchop: brings in the photographed saree, shirt, bridal piece and masks, with a little Hinglish. Adds red silk/handloom assumptions, a physical visit promise without an address, and “we answer every call ourselves.” The old version also invents hours, custom sizing and colourfastness.
6. r23-native-website-generation-ca4c7db8-1785-4909-acb1-24b8a1b4ff4f-round2 — Skanda Tex: retains supplied 1800 DPI, 920 m²/H and printer technology, and is much clearer than the baseline’s exaggerated technical claims. Still treats draft design samples as completed production and describes an actual workflow/press-operator contact not established by the source. No independent print or site rendering verification.

### Native website editing

**Hold.** Simple operations apply; full intent and copy consistency remain unproven

Three operation plans apply and prepare successfully; three requests return clarification. The rename is broadly sensible, but the address edit updates a structured contact field while old-location editorial copy remains. The maps flow leaves an earlier requested address correction unresolved. Historical finalized documents do not reliably expose the original provider operation result.

**Review scope:** All six complete edit plans, user turns, supplied conversation messages and targeted base copy/contact/SEO fields were read directly. Actual application functions applied three plans and validated prepared documents. Full website markup and visual rendering were not manually reviewed.

**Current chain:** gemini-3.7-flash → gpt-5.6-luna  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-luna  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r24-native-website-edit-6dafbcb1-b41f-4ccc-b63c-c5466dc2670b-round2 — Store terminology: updates category, relevant labels and SEO with narrow operations that apply. One setCopy is an exact no-op; the hero wording changes slightly beyond a literal replacement. Existing unverified merchandising claims are inherited. Broadly reasonable narrow edit.
2. r24-native-website-edit-60774f29-58bd-40fe-8943-bf84908cf2c1-round2 — Phone-only input: asks which contact field and formatting to use. With no conversation supplied this is defensible, though it adds friction to an obvious likely phone replacement. It does not complete the change; cannot count clarification as a successful edit.
3. r24-native-website-edit-4f99e874-0fd9-4ba6-8e48-b14645d81532-round2 — More photos and premium style: correctly notices only the existing photo is available and asks for assets. The source requires all-or-nothing handling of compound requests, explaining why restyling is deferred. No evidence of improved visual quality.
4. r24-native-website-edit-93bc889e-c818-47e1-97c6-416d5b3a606b-round2 — New address: sets the exact supplied line and applies successfully. Existing hero copy still says Kuchha Kathara, and coordinates/maps are preserved. A field update alone does not prove the site displays a consistent corrected location.
5. r24-native-website-edit-5e2738ef-8e4d-412e-8fc6-43e15ba17e54-round2 — Maps URL: updates the exact URL successfully using conversation context. The same history includes a request to remove the wrong locality; no operation resolves that copy/address issue. Stored finalized baseline contact state also appears unchanged, so it is not a clean operation-level comparator.
6. r24-native-website-edit-bb8189ec-cddf-4c8f-93ba-3c8f2a19f164-round2 — Maps after remove-address request: clarification acknowledges both intents and avoids guessing. Reasonable under inconsistent base/history, but the requested change remains incomplete and is not evidence of replacement readiness.

### Website builder prompt

**Hold.** Detailed build instructions, but serious fabricated-business and placeholder-contact content.

All six preserve the main theme colours and image-URL rules. Missing historical Brand Positioning fields limit the comparison, yet DeepSeek independently instructs builders to add invented services, hours, phone numbers or fake form success. These are unacceptable default behaviors even with sparse inputs.

**Review scope:** Direct reading of all six reconstructed requests and paired prompt text. Appended baseline asset URL lists were not treated as generated text; full historical positioning, asset context and requested languages were not reconstructed. No website was built from these prompts.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Source comments: Flash-Lite defaults to zero thinking → reasoning.enabled=false

1. r25-website-builder-prompt-283dc549-ca73-4355-88d5-b338fa9a2bf6-none-round2 — Invents a phone number, natural/cruelty-free/small-batch claims and skincare consultations from a sparse category/name. Historical product positioning was absent from reconstruction, so tagline differences are not fair model failures.
2. r25-website-builder-prompt-db6f66c0-68d8-44dc-9dc8-4126035624f3-none-round2 — Preserves real phone/address but invents emergency plumbing services and opening hours. Allows footer logo despite header-only constraint; input omits historical pump catalogue and Hinglish context.
3. r25-website-builder-prompt-2e3fbd58-2317-46fd-a606-806122a90527-none-round2 — Correctly omits missing phone/address, but asks a nonfunctional form to show success and suggests unverified financial offerings. Footer logo contradicts header-only rule; missing historical language context limits comparison.
4. r25-website-builder-prompt-f5a2fd16-f938-49a2-a99a-082b0a6ff0e9-none-round2 — Correct phone/palette. Adds same-day delivery, genuine products, doorstep support and daily hours without evidence; historical Hindi/USP inputs are absent.
5. r25-website-builder-prompt-bfe7716a-8d34-4da4-9ce1-ed9c713a3a31-none-round2 — Builds a baked-goods shop rather than the saved commercial-oven business because reconstruction lacks product positioning. Also invents opening hours, custom cakes and catering; cannot approve from this pair.
6. r25-website-builder-prompt-4a8586eb-a0e1-4ce2-9875-1f6fcbbdd00e-none-round2 — Preserves location and avoids a made-up phone number, but leaves a nonworking tel link, placeholder socials and unsupplied stats/operational promises. Requires real contact/data completion.

### Website theme suggestions

**Candidate.** Comparable, with more distinct bold options

All six sets offer a credible minimal, bold and professional direction. DeepSeek often separates the bold palette more clearly than the saved sets. These are design proposals, not verified brand facts or rendered/contrast-tested themes.

**Review scope:** Direct reading of captured inputs and both outputs; generated media is not rendered.

**Current chain:** gemini-3.5-flash-lite  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.5-flash-lite  
**Thinking tested:** disabled. Source comments: Flash-Lite defaults to zero thinking → reasoning.enabled=false

1. r25-website-themes-283dc549-ca73-4355-88d5-b338fa9a2bf6-none-round2 — Distinct botanical, purple/orange and navy/gold directions; bolder than the saved mostly muted palettes.
2. r25-website-themes-db6f66c0-68d8-44dc-9dc8-4126035624f3-none-round2 — Good water/industrial/copper differentiation. Emergency badges are only a proposed design treatment and need real service availability before use.
3. r25-website-themes-2442a6c5-acda-4467-b674-b26bfbf161ac-none-round2 — Useful separation of quiet ivory, Indian-festival color and professional navy.
4. r25-website-themes-2e3fbd58-2317-46fd-a606-806122a90527-none-round2 — Plausible investment styling but claims inspiration from photos that were not supplied; present them as proposed directions only.
5. r25-website-themes-f5a2fd16-f938-49a2-a99a-082b0a6ff0e9-none-round2 — More differentiated light/dark and warm/cool choices than the historical all-dark themes. Instant-delivery wording should not become factual website copy.
6. r25-website-themes-bfe7716a-8d34-4da4-9ce1-ed9c713a3a31-none-round2 — Coherent minimal, playful and premium bakery directions; generated layout and accessibility were not evaluated.

### Click-to-WhatsApp ad copy

**Hold.** Mixed; existing is more cautious about missing business facts

The messages are generally usable, but DeepSeek asserts online-only/pan-India service when the profile merely lacks a location. That can mislead owners about targeting rationale. Sparse profiles need a safer uncertainty policy.

**Review scope:** Direct reading of captured inputs and both outputs; generated media is not rendered.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r27-ad-ctwa-1fb0d6c3-c080-4711-908f-34608db8345d-fresh — Inkly60: good Hinglish sales copy and correct Rs. 349 price, but online-only operation is invented. Prefer existing rationale until this is fixed.
2. r27-ad-ctwa-57bdbd78-0dea-4053-92bc-42ebb9e55c10-fresh — Auspora: sensible local cosmetics recommendation. Wider radius and all-gender targeting are reasonable choices, not proven improvements; Hinglish language is not specified by the input.
3. r27-ad-ctwa-5c6dd76a-f6c1-4a95-99f4-7540f67b4722-fresh — KPD: both fill a very sparse profile with generic latest/new apparel copy. DeepSeek explains missing location more honestly here; neither establishes national fulfillment.
4. r27-ad-ctwa-1c9d2483-4ef4-4893-bc8b-f703da20a589-fresh — Garam Masala: concrete local Chinese-food copy and a useful menu/timing inquiry. Both usable; radius/age differences need campaign evidence, not text preference.
5. r27-ad-ctwa-07c8c5bd-2b0f-4cbb-ad0a-9fa5e04b29b1-fresh — Promotion: DeepSeek invents an online business that can serve anywhere in India from essentially empty facts. Hold this behavior.
6. r27-ad-ctwa-a02c732d-020e-4e6e-ba95-fcef458d8d16-fresh — Children clinic: clear appointment-focused copy without treatment guarantees. Both usable; no evidence that one targeting band performs better.

### Engagement ad copy

**Hold.** DeepSeek often writes a stronger engagement prompt, but grounds facts less reliably

DeepSeek more consistently asks a natural question suited to comments, while several existing outputs read like visit-the-store ads. However, two DeepSeek rationales invent pan-India shipping, which outweighs the creative advantage for unrestricted rollout.

**Review scope:** Direct reading of captured inputs and both outputs; generated media is not rendered.

**Current chain:** gemini-3.7-flash → gpt-5.6-terra  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash → gpt-5.6-terra  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r27-ad-engagement-1fb0d6c3-c080-4711-908f-34608db8345d-fresh — Inkly60: the style-choice question better fits engagement, but the rationale falsely treats pan-India shipping as supplied fact.
2. r27-ad-engagement-57bdbd78-0dea-4053-92bc-42ebb9e55c10-fresh — Auspora: asking for a favorite beauty product is more directly engaging than the baseline store invitation. Both use a defensible local audience.
3. r27-ad-engagement-5c6dd76a-f6c1-4a95-99f4-7540f67b4722-fresh — KPD: both invite style opinions; DeepSeek adds every-occasion positioning to a sparse profile. No clear overall winner.
4. r27-ad-engagement-1c9d2483-4ef4-4893-bc8b-f703da20a589-fresh — Garam Masala: local favorite-dish/friend question is more conversational and engagement-oriented than the baseline visit CTA.
5. r27-ad-engagement-07c8c5bd-2b0f-4cbb-ad0a-9fa5e04b29b1-fresh — Promotion: generic copy is unavoidable with this input, but DeepSeek must not assert that the business ships nationally.
6. r27-ad-engagement-a02c732d-020e-4e6e-ba95-fcef458d8d16-fresh — Children clinic: the parent-routine question encourages relevant comments without diagnosis or promised outcomes. Stronger fit to engagement than the baseline clinic introduction.

### Google post polishing

**Hold.** Mostly equivalent; one unsupported language label

Text preservation is generally good, but KPD is labelled Polish without evidence. Five of the six drafts are only names or near-empty fragments, so this is weak evidence for real merchant-draft polishing.

**Review scope:** Direct reading of captured inputs and both outputs; generated media is not rendered.

**Current chain:** gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r28-google-post-polish-1fb0d6c3-c080-4711-908f-34608db8345d-fresh — Inkly60: both preserve the slogan and starting price; punctuation changes are appropriate. en-IN versus en is acceptable.
2. r28-google-post-polish-57bdbd78-0dea-4053-92bc-42ebb9e55c10-fresh — Auspora: DeepSeek conservatively preserves the name; Gemini adds Welcome. Both are harmless but the input does not test substantive polishing.
3. r28-google-post-polish-5c6dd76a-f6c1-4a95-99f4-7540f67b4722-fresh — KPD: text is preserved, but DeepSeek returns pl. A three-letter name cannot establish Polish; use known merchant locale or explicit uncertainty.
4. r28-google-post-polish-1c9d2483-4ef4-4893-bc8b-f703da20a589-fresh — Garam Masala: both preserve the restaurant name; no meaningful polishing challenge.
5. r28-google-post-polish-07c8c5bd-2b0f-4cbb-ad0a-9fa5e04b29b1-fresh — Promotion: both remove redundant punctuation and preserve the word; no meaningful content challenge.
6. r28-google-post-polish-a02c732d-020e-4e6e-ba95-fcef458d8d16-fresh — Children clinic: both preserve meaning; Gemini improves punctuation/capitalization slightly more. No invented services.

### Google post writing

**Hold.** Mixed; better grounding in some cases

The visual descriptions are generally accurate, but a profile/photo conflict produces unusable hybrid copy, the storefront case adds unsupported newness and questionable OCR, and the restaurant case invents signature status. Both models need stronger handling of ambiguous inputs.

**Review scope:** Direct reading of all six inputs and both outputs, plus visual inspection of all six source photos. No generated post images evaluated.

**Current chain:** gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r28-google-post-write-1fb0d6c3-c080-4711-908f-34608db8345d-fresh — Accurately describes the navy adventure tee and visible slogan; useful copy without a specific unsupported fabric claim.
2. r28-google-post-write-57bdbd78-0dea-4053-92bc-42ebb9e55c10-fresh — Both displayed phone numbers match the sign, but Dr.Santhiyas is questionable OCR and storefront now features invents newness. Too much sign transcription for a customer update.
3. r28-google-post-write-5c6dd76a-f6c1-4a95-99f4-7540f67b4722-fresh — Apparel profile conflicts with a café photo. Existing output invents a café business; DeepSeek combines café imagery with an apparel pitch. Neither is ready to publish.
4. r28-google-post-write-1c9d2483-4ef4-4893-bc8b-f703da20a589-fresh — Visible curry description is plausible, but our signature curry adds unprovided status; lunch/dinner availability is also assumed. Existing output makes similar availability assumptions.
5. r28-google-post-write-07c8c5bd-2b0f-4cbb-ad0a-9fa5e04b29b1-fresh — Both models produce useful outfit copy grounded in the embroidered kurta and white trousers; maroon versus magenta is a minor lighting interpretation.
6. r28-google-post-write-a02c732d-020e-4e6e-ba95-fcef458d8d16-fresh — DeepSeek accurately describes the clinic poster and entrance without inventing treatments. Repeating an undated health poster is less useful as a current update and needs owner context.

### AI video editing director

**Hold.** Cannot compare creative quality against saved baselines: those contain aggregate render metrics, not director JSON.

Five outputs and one exhausted connection failure. Returned directions invent face/product positions without a frame, and some timing conflicts are visible directly in the output. This service needs complete director baselines, frames and render validation.

**Review scope:** Direct reading of all six transcripts, five final outputs and one failure record. Baselines contain metrics only; no representative frames or rendered edits were available.

**Current chain:** gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-3.7-flash  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r29-ai-director-81aba38e-7359-4653-aa41-c1646bdcf3f6-round2 — Useful transcript-specific emphasis, but a cancelled-flight insert changes a visa-rejection story and frame positions are asserted without an image. B-roll gaps exceed the requested spacing.
2. r29-ai-director-0d8613cd-6fae-4eb3-8965-dfd5f39a3fed-round2 — A 1.5-second insert begins at 62% of a 2.32-second clip and overruns its end. The short clip cannot satisfy all template timing rules; the model should adapt more carefully.
3. r29-ai-director-accb21c0-d966-4d45-98ca-9b3b11d64876-round2 — Relevant spoon/shop inserts. Safe-zone explanation says hands occupy center-low but puts captions center-lower there; no frame supports any placement.
4. r29-ai-director-3402493c-5d7d-424d-a8f8-29c0cf6e3d8e-round2 — Four inserts are too tightly spaced for the requested three-second minimum and some overlap. Product and CTA positions are guessed.
5. r29-ai-director-f8c1155f-d5c6-481d-b72d-a25fd7fca9a6-round2 — Treats the spoken production brief as ad content. Generic restaurant visuals are plausible, but timing and on-screen placements are not verified; no full historical JSON exists.
6. r29-ai-director-4828d071-78ba-4a27-83c5-a00030154df7-round2 — No final output after five connection-error attempts. Retain as a failed case, not a missing or successful sample.

### Avatar gender labels

**Limited candidate.** Equivalent on these inputs

All six prompts explicitly describe a man and both models return male. This is straightforward extraction, with no observed disagreement; the sample says nothing about female or ambiguous prompts.

**Review scope:** Direct reading of captured inputs and both outputs; generated media is not rendered.

**Current chain:** gemini-2.5-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gemini-2.5-flash  
**Thinking tested:** high. Default/automatic → high approximation; must be calibrated for latency and output budget.

1. r31-avatar-gender-23c0e251-3223-4310-868d-3c529d876824-round2 — Correct male label for an explicitly described adult Indian man.
2. r31-avatar-gender-ab9cb40f-7106-4496-9df2-c7eb346cb7bf-round2 — Correct male label for an explicitly described middle-aged Indian man.
3. r31-avatar-gender-4bff8476-f246-4402-928d-8beaba8aa22c-round2 — Correct male label for an explicitly described man in a navy suit.
4. r31-avatar-gender-bbe1d475-49c7-4cc5-abfe-21abc4a9fc71-round2 — Correct male label for an explicitly described man at a riverside.
5. r31-avatar-gender-4f1ddacf-0ab0-4875-b3ab-6dfd50e760da-round2 — Correct male label for an explicitly described middle-aged man in a blazer.
6. r31-avatar-gender-c343487f-7b10-4e7b-8e51-3629227d4e78-round2 — Correct male label for an explicitly described Indian man in a windbreaker.

### Garment analysis

**Hold.** Mostly plausible; outfit/background boundary is inconsistent

The main garment readings are mostly convincing. A visible background shirt is included as outerwear in one case but ignored in similar neighboring images; this could add unwanted clothing to generation. The red dress disagreement is plausibly a historical label problem, not evidence DeepSeek is wrong. No saree cases were sampled.

**Review scope:** Direct reading of all six inputs and both outputs, plus visual inspection of all six source images. Missing historical fields such as isSaree were not treated as disagreements.

**Current chain:** gpt-5.6-luna → gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gpt-5.6-luna → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r34-garment-analysis-dbaf806f-7d9f-4828-9fbb-16c1b18206a1-round2 — Correctly inventories the printed shirt plus separate brown trousers as a full set.
2. r34-garment-analysis-fcb616be-00f7-4cd4-8c94-483bf9db50c5-round2 — Correctly identifies the cream pleated wrap dress as one-piece/full.
3. r34-garment-analysis-fc572d66-1f8d-417e-a8b7-4d0f36f477b0-round2 — Correct one-piece/full identification for pink embroidered tiered dress.
4. r34-garment-analysis-beb93713-ffde-448d-8882-bc20d2e11c28-round2 — The olive shirt is visibly on the rack behind the ivory dress, not established as its outerwear. Including it risks recreating background stock as the outfit.
5. r34-garment-analysis-e223c234-df88-4f2d-9151-fd6cf7a29734-round2 — Correct brown tiered dress label; ignores the same background shirt included in the previous case, showing inconsistent segmentation.
6. r34-garment-analysis-17375b50-23e0-4511-8180-4c5b8d150ae7-round2 — One-piece/full is plausible for the long button-front tiered embroidered dress and looks better than saved upper-only. Owner category remains the final authority.

### Garment prompt writing

**Hold.** Generally comparable and better at explicitly retaining no-pants requests, with a face-identity omission.

Framing and background instructions are consistent across six outputs, and user no-pants overrides are preserved. One output drops the explicit face-reference instruction. All cases refer to an additional face image but the captured requests contain only the garment image, limiting identity evaluation.

**Review scope:** Direct reading of all six prompts and both outputs; source garment images had been visually checked during garment analysis. No generated images or separate face references were available.

**Current chain:** gpt-5.6-luna → gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gpt-5.6-luna → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r34-garment-prompt-dbaf806f-7d9f-4828-9fbb-16c1b18206a1-round2 — Preserves two-piece outfit, collar detail, alley and full front view. Generic fair-complexion text can conflict with the requested face identity; actual face reference is missing from captured input.
2. r34-garment-prompt-fcb616be-00f7-4cd4-8c94-483bf9db50c5-round2 — Clearly honors no pants, full framing, grey background and requested accessories. More explicit user-instruction retention than baseline.
3. r34-garment-prompt-fc572d66-1f8d-417e-a8b7-4d0f36f477b0-round2 — Uses frock because the user explicitly names it, an allowed exception. Keeps no pants and accessories; one-hand-at-hair pose makes arms-at-sides wording slightly inconsistent.
4. r34-garment-prompt-beb93713-ffde-448d-8882-bc20d2e11c28-round2 — Keeps the ivory outfit and omits background stock garments correctly, but completely drops the requested reference face identity.
5. r34-garment-prompt-e223c234-df88-4f2d-9151-fd6cf7a29734-round2 — Keeps no lower garment and includes accessories, unlike the baseline’s no-accessories instruction despite the user request.
6. r34-garment-prompt-17375b50-23e0-4511-8180-4c5b8d150ae7-round2 — Correctly prioritizes no pants over the contradictory inferred missing-lower-garment default. Preserves full front framing and face-reference instruction.

### Product analysis

**Hold.** Better labels and specificity; some shots invent unseen construction

DeepSeek supplies useful leaf labels, including barbell where the saved label says dumbbell, and specific shot concepts. But it proposes a wooden-door installation for what appears to be glass-door hardware and invents an unseen jacket back. Classification alone is more promising than the combined shot-idea output.

**Review scope:** Direct reading of all six inputs and both outputs, plus visual inspection of all six source images. Resulting generated photos were not produced.

**Current chain:** gpt-5.6-luna → gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gpt-5.6-luna → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r35-product-analysis-316cb796-c253-40e3-9c1c-4a8ef125fb1a-round2 — Correct broad lock label and detail ideas. The pictured clamp-style glass hardware does not establish wooden front-door compatibility.
2. r35-product-analysis-542126fb-6d94-4663-bc3b-34e67cc41dd5-round2 — Correct flour-mill label and coherent range of concepts. Multiple models are shown, so the kitchen-counter scale and specific material/performance need confirmation.
3. r35-product-analysis-5c87e2ac-2068-47a1-9319-e96ee058fe92-round2 — Footwear label and brown sandals/Soft Cushion/City Foots concepts are grounded and much better than the historical generic fallback ideas.
4. r35-product-analysis-8575c3ee-49bf-498d-adbc-63193ad3c9df-round2 — Correct bomber-jacket label and useful trim details, but the described color-block back is not visible in the front-only image.
5. r35-product-analysis-8ab01e2e-38bc-405d-a0fc-7a7647bfc228-round2 — Barbell is more accurate than historical dumbbell; five shots are complementary and appropriate.
6. r35-product-analysis-ab09d96d-61d6-4288-a2d8-01d7b57e182e-round2 — Correct motorcycle label and plausible engine/studio/workshop/scale shots; cosmetic finish detail is less certain at this image resolution.

### Product prompt writing

**Hold.** Specific in simple cases, weak on constrained variants

The output-key conflict and HEIC transport are fixed in the eval adapter. Direct quality review still finds invented fabric properties, an insufficient no-face instruction and a generic handheld treatment of a jacket. Formatting success is not enough to approve this service.

**Review scope:** Direct reading of all six inputs and both outputs; supplied product and avatar images visually inspected (the repeated sandal source was inspected with product analysis). No downstream images/videos generated.

**Current chain:** gpt-5.6-luna → gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gpt-5.6-luna → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r35-product-prompt-d4d2890d-4cdf-4f90-a4ca-7ab263e0e537-round2 — Good bird-spike presentation and consistent starting composition; product details match the image. The person/woman wording conflict is in the source prompt.
2. r35-product-prompt-5312edab-23ea-48a9-8a62-584deff2f767-round2 — Correct pickle jar and outdoor setting. At chest height on a table is spatially awkward; needs clearer pose wording.
3. r35-product-prompt-99e85544-7f1a-44ce-a736-3a25fd8d9699-round2 — Suitable studio sandal prompt with true-color/detail preservation; video combines push-in and light sweep rather than one motion.
4. r35-product-prompt-57bf432b-0ebc-4758-9ab5-bc0973bb5357-round2 — Reasonable outdoor sandal setup; the video stacks a push-in, light sweep and buckle adjustment into five seconds.
5. r35-product-prompt-1dcad26a-0f73-4adf-85a1-eb143f45c85f-round2 — The input product is a blue patterned jacket, but output only says product and has the woman holding it in a kitchen. It misses the actual wearable-use brief.
6. r35-product-prompt-3a7fee5c-726f-4dde-90d2-295ff713465c-round2 — Invents lightweight/breathable callouts, and without a dominating face does not mean no face. The source also contradicts its own callout request with a no-text rule; resolve that before fair approval.

### Avatar promotional script

**Hold.** More conversational, but invents testimonial events

The copy often sounds natural, and visible product text is used well. However, it fabricates personal/customer experiences and stronger exclusivity than supplied. Historical scripts sometimes include user edits and mismatched current business context, so this is not a clean model-only win-rate comparison.

**Review scope:** Direct reading of six inputs and both outputs plus the three supplied images. No audio generated; duration is judged as script suitability only.

**Current chain:** gpt-5.6-luna → gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gpt-5.6-luna → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r36-avatar-script-81aba38e-7359-4653-aa41-c1646bdcf3f6-round2 — Invents the creator having used the visa service successfully. Current gift-shop profile conflicts with the travel brief, making the historical comparison unreliable.
2. r36-avatar-script-0d8613cd-6fae-4eb3-8965-dfd5f39a3fed-round2 — Ganesh/Om, 240 GSM, cotton, oversized and Sai Home Mart are all visible in the supplied poster; the energetic script is grounded.
3. r36-avatar-script-2ed5a70c-7eb2-4343-a6fe-f48c76e20129-round2 — No one else has overstates exclusive gift/art positioning, and the new DM keyword workflow is not supplied. Historical copy is more restrained.
4. r36-avatar-script-21b20488-29a4-4372-ae52-e4c6f6343e1e-round2 — Invents yesterday’s conversation with an aunty and merchant quotes. This is not genuine social proof supplied in the brief.
5. r36-avatar-script-accb21c0-d966-4d45-98ca-9b3b11d64876-round2 — Reasonable utensils copy grounded in supplied range/price/retail-wholesale facts, though longer and more slogan-heavy than the baseline.
6. r36-avatar-script-3402493c-5d7d-424d-a8f8-29c0cf6e3d8e-round2 — Clear local gaming CTA and correct supplied console range; similar usefulness to saved copy, subject to actual spoken timing.

### Voice script processing

**Hold.** Five acceptable cases; one clear conversion failure

The preservation instruction improves the previous failure but does not fix it: the English production brief is still translated to Hinglish despite an explicit instruction to preserve English. This service should not migrate yet.

**Review scope:** Full source-script segments and complete outputs read directly. TTS pronunciation and audio quality are not evaluated.

**Current chain:** gpt-5.6-luna → gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gpt-5.6-luna → gemini-3.7-flash  
**Thinking tested:** low. low, matched to current primary setting; existing fallback keeps its own setting.

1. r37-voice-script-81aba38e-7359-4653-aa41-c1646bdcf3f6-round2 — Converts Hindi words while preserving the travel-company name, 35+ years, website, dramatization note and visa disclaimer. Close to the saved version.
2. r37-voice-script-0d8613cd-6fae-4eb3-8965-dfd5f39a3fed-round2 — Preserves the English t-shirt sentence unchanged. This follows the conversion rule more faithfully than the saved translated/altered version.
3. r37-voice-script-accb21c0-d966-4d45-98ca-9b3b11d64876-round2 — Preserves the spoon advertisement, retail/wholesale facts, seller/location and expression tags; differences are punctuation and colloquial Hindi spelling.
4. r37-voice-script-3402493c-5d7d-424d-a8f8-29c0cf6e3d8e-round2 — Preserves gaming-console names, shop name, CTA and curious tag; effectively the same as saved output.
5. r37-voice-script-d8510d85-2820-490d-b1de-798eaa41e823-round2 — Preserves the therapy service list, location, slogan and serious tag without adding treatment claims; punctuation differs.
6. r37-voice-script-f8c1155f-d5c6-481d-b72d-a25fd7fca9a6-round2 — Fails the English-preservation boundary: translates the English production brief to mixed Hindi/English. Retaining the original number and CTA does not make this compliant.

### Sora prompt enhancement

**Hold.** More detailed direction, but worse brevity, timing and some user-intent preservation.

All six return usable text structures. The tuition script is far too long for ten seconds, the five-second tee ad remains unspeakably dense, and the Pichola request loses its continuous full-length framing. The Mediterranean ad adds a pure-linen claim and changes the actor’s attire despite preservation instructions.

**Review scope:** Direct reading of all six source requests, shared meta-prompt and both final texts. Image pixels and generated video were not independently evaluated for this service; image/material claims remain unverified.

**Current chain:** gpt-5.6-luna → gemini-3.7-flash  
**Proposed after approval:** deepseek/deepseek-v4.1-flash → gpt-5.6-luna → gemini-3.7-flash  
**Thinking tested:** high. OpenAI medium → DeepSeek high

1. r38-sora-enhancement-v2-0f3df58e-6565-40e7-9183-5b1d6468f52b-round2 — Preserves location/admissions intent but puts roughly a full paragraph of speech into ten seconds. A detailed address plus three other lines cannot be delivered naturally at the requested pace; baseline is also overlong.
2. r38-sora-enhancement-v2-0fdecddb-0b08-404f-876d-bac7fa577c38-round2 — Keeps location, lens and actor reference, but inserts mid/detail cuts where the user specified a backward dolly keeping the model head-to-toe. Existing output stays closer to the camera request.
3. r38-sora-enhancement-v2-f73432ea-5527-4fe4-b3ad-55fd857c5195-round2 — Preserves the low-to-high camera move and fort setting. Adds linen material and a personal-experience voiceover; source pixels were not independently checked in this review.
4. r38-sora-enhancement-v2-c81167e6-9908-411f-866f-56ad6a8a1250-round2 — Shorter than the saved baseline but still about thirty spoken words in five seconds, plus multiple visual cuts. Both require a duration or copy tradeoff.
5. r38-sora-enhancement-v2-3aea1632-c0d5-4fae-ba47-30a127896255-round2 — Honors no voiceover/no speech throughout and keeps the outfit central. More detailed cutaways, comparable creative usefulness; image identity was not rendered/tested.
6. r38-sora-enhancement-v2-0c911531-b2c3-43cd-a913-1fcac70604be-round2 — Plausible villa setting, but asserts pure linen, switches to Hinglish and dresses the actor in the product although the meta-brief says preserve actor attire. Existing output uses a hanger to preserve attire.

## Inspect and reproduce

[Interactive comparisons](review.html) contain the full captured prompts and final outputs beside each direct judgment. [Summary JSON](summary.json) contains mechanical totals, [written reviews](llm-reviews.json) contain the judgments, and [source manifest](source-manifest.json) records hashes. `cases.json`, `results/` and `baselines/` retain reproducible inputs and responses. Do not rerun earlier case builders over the final manifest. Changed requests require new case IDs.

No application source, model routing, staging content or deployment was changed. This report is publicly hosted at the owner’s request. [Full Gemini inventory](gemini-text-thinking-inventory-2026-09-10.md) and [full OpenAI inventory](openai-text-reasoning-inventory-2026-09-10.md) include versions beyond the selected replacement scope.
