Skip to content

Choose better models,
with evidence.

Run controlled, blind evaluations across generative media. Generate on fal or bring finished outputs, then turn team judgment into a decision you can defend.

  • Image
  • Video
  • Audio
  • Text
  • Multimodal
Material fidelity evaluation
Product preview · Blind pairwise
A chrome ring threaded with peach silk in a warm controlled studio sceneCandidate A

Identity hidden

An alternate composition of the same chrome ring, peach silk, and studio materialsCandidate B

Order varies by reviewer

Prefer ATiePrefer B
Same prompt. Same seed. Model identity stays hidden until the evaluation closes.

01 Controlled inputs

02 Media-native review

03 Decision-ready evidence

One evaluation system

Everything around the vote matters.

GenMedia keeps the question, generation contract, blind presentation, reviewer context, and statistical evidence in one traceable workflow.

Release candidate evaluationEvidence building
A sequence of paper and graphite forms threaded by a green invariant

Current evidence

Candidate B leads

Directional signal across motion, temporal consistency, and prompt adherence.

Candidate A38%
Candidate B62%
24 prompts18 reviewers92% coverage
01

Controlled generation

The endpoint schema is part of the Evaluation Task

GenMedia resolves current fal OpenAPI contracts, maps each prompt or reference input explicitly, and preserves the exact run lineage behind every output.

02

Media-native review

Judging tools match the artifact

Compare images, video, audio, text, and 3D-oriented outputs with blind pairwise or rating workflows. Video pairs can move into DeltaFrame for frame-level inspection.

03

Statistical context

A ranking is not automatically a conclusion

Weighted rankings sit beside uncertainty intervals, pair coverage, multiplicity-aware tests, position checks, and clear language when the evidence is incomplete.

04

Traceable evaluations

Prompts and decisions stay connected

Prompt packs, projects, taxonomy tags, immutable runs, and privacy-safe exports preserve the path from an evaluation to a release decision.

One trail from Evaluation Task design to release decision.

A deliberate operating loop

From a question to a confident release.

  1. 01

    Define the decision

    Choose the comparison intent, media modality, held-constant inputs, audience, and success criteria before any candidate is reviewed.

  2. 02

    Bring or generate media

    Import finished outputs from any source or generate directly from schema-verified fal endpoints with explicit prompt and input bindings.

  3. 03

    Review without the reveal

    Randomized candidate order, completion gates, and media-specific controls keep model identity out of the judgment loop.

  4. 04

    Act on measured evidence

    Read the fast summary first, then inspect uncertainty, pair coverage, prompt effects, bias checks, and the next best sampling action.

DeltaFrame

See the difference,
frame by frame.

Step through synchronized video, compare split and difference views, and jump to the exact moments where motion or detail breaks down.

Explore DeltaFrame
A · 00:08.42
B · 00:08.42

Evidence that respects people

Good evaluation protects the reviewer, too.

Blind protocols reduce bias. Uncertainty stays visible. Reviewer reliability is handled conservatively and kept out of ordinary result views.

01

Blind by default

Evaluator-facing candidate order is deterministic and shuffled. Model identities and aggregate results remain hidden until the task completion gate is satisfied.

02

Ambiguity is protected

Reviewer reliability uses independent peer pools and treats strong cross-pool disagreement as uncertain evidence, not as an automatic reviewer failure.

03

History is append-only

Assigned inputs and past outputs are not silently rewritten. New prompts or endpoints create a new run while earlier lineage remains inspectable.

Ready when the question is

Build the eval before
the conclusion.

Create a controlled Evaluation Task in the workspace, or use the documentation to align your team on protocols, roles, and evidence.