When inference cost or latency is blocking production.
Here is my model and serving setup: [paste]. Propose optimizations: quantization, distillation, batching, caching. Rank by expected speedup vs effort and accuracy risk. Flag the trade-off of each.
Anything in square brackets is meant to be replaced. The builder turns those into fields you can type into.
Tired of re-pasting the same context into every new chat? That is what the Unimatrix MCP server is for.