307 USA Computing Olympiad problems with official unit tests, reference solutions and analyses, used to test algorithmic reasoning via inference methods including self-reflection, retrieval, and human-in-the-loop hinting.
discriminates
[build-verify-reflect][teacher-student]
saturation
open
patternLabFit
The paper's own best inference method combines self-reflection with retrieval over episodic knowledge, a direct build-verify-reflect precedent. Its human-in-the-loop hint study, where targeted hints let GPT-4 solve 13 of 15 previously-unsolved problems, is direct empirical evidence for a teacher/hint-giver role improving a solver role, i.e. teacher-student.
Distinct from livecodebench (fixed problem set, not contamination-refreshed) and codecontests (proposed alongside it, a training/eval-scale dataset rather than a fixed olympiad set): USACO's defining feature is the difficulty-tiered structure plus an explicit human-in-the-loop hinting study, giving direct empirical evidence for a teacher/hint-giver pattern rather than pure self-consistency. Tier-name claim ('bronze/silver/gold/platinum') came from a third-party summary, not the fetched abstract, so it is omitted here.