Tag: Long-Context LLMs

Sep 21
Needle In A Haystack (NIAH) for Agents: Locating Ephemeral Instructions in 1M+ Token Contexts

In passive document question-answering, the Needle In A Haystack (NIAH) benchmark was designed to measure whether a foundation model could retrieve a single factual sentence placed at varying depths within a massive context window. An arbitrary fact—such as stating that a specific pizza topping is preferred in a fictional city—is inserted at an arbitrary depth […]