How long can your skills be before your agent forgets what you told it? - Laurie Voss, Arize AI

AI Engineer22:25
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Laurie Voss revisits IFScale, a benchmark that asks models to include specified words in a business report and scores how many appear. After reproducing earlier degradation with older models, he extends the vocabulary beyond the original test because newer models perform strongly at its previous upper bound. He reports substantially greater capacity, with results varying by model.

    Laurie Voss describes distinct failure patterns in his experiments: DeepSeek V4 Pro gradually omits words, Claude Opus 4.7 can refuse requests containing concerning word combinations, Gemini 3.1 Pro can exhaust its output budget through extensive reasoning, and GPT-5.5 can begin a report before declining to complete it. He explains that filtering the random vocabulary affected the Claude testing, a methodological detail relevant to interpreting comparisons.

    Laurie Voss reports roughly 99 percent keyword inclusion for GPT-5.5 at 5,000 constraints in this test, but explicitly distinguishes keyword tracking from following realistic instructions. The benchmark does not establish clear reasoning across a long prompt, resolution of conflicting requirements or robustness to changes in wording and order. He discusses related research as reasons to retain those caveats.

    Laurie Voss argues that engineers should revisit old assumptions about keeping every skill extremely short while still weighing cost and latency. His practical emphasis is verification: polished partial output can conceal failure, so applications need checks against their actual requirements rather than treating benchmark capacity as a guarantee of reliable completion.

    Original YouTube thumbnailWatch on YouTube