What is system-level artificial intelligence evaluation?
Definition
System-level artificial intelligence evaluation tests how models, agents, tools, memory, controls and human workflows operate together. Measures can include component accuracy, task completion, user frustration, concurrency, cost, security and recovery from failures.
The approach catches problems that isolated model scores miss, such as invalid handoffs or agents bypassing restrictions through alternate tools. Results should be tied to representative workflows and the consequences of failure.
Acronyms and aliases
end-to-end AI system evaluation synonymsystem-level AI evaluation variant
General terms
Related terms
Frequently asked questions
How is system-level AI evaluation different from model evaluation?
It evaluates the complete application and workflow, including tools, state, people and controls, rather than one model response.
What should system-level AI evaluation measure?
It should measure task outcomes, component quality, user effort, concurrency, cost, security boundaries and failure recovery.