Skip to main content
AI / LLM PromptingCurated

Evaluation Set Generator

When you have a model but no way to tell if a prompt change made it better or worse.

The prompt
I need to evaluate an LLM on [task]. Generate 20 diverse test inputs covering edge cases, ambiguous inputs, and adversarial phrasings. For each, write the criteria a correct answer must meet (not the answer itself).

Anything in square brackets is meant to be replaced. The builder turns those into fields you can type into.

Tired of re-pasting the same context into every new chat? That is what the Unimatrix MCP server is for.

Browse the rest of the library