The longer-term goal is to reconstruct the task × model × harness success matrix without running every cell. I will test whether success has enough low-rank structure for SVD to identify a small set of Terminal-Bench tasks that predicts the full matrix.
If that works, the reduced task panel becomes a fast screen for many more model–harness combinations which leave space to reserve full runs for combinations that look promising.
The first run held the model fixed—DeepSeek V4 Flash, high reasoning—and compared six popular harnesses over roughly 1,500 Terminal-Bench rollouts.