N
Hacker Next
new
past
show
ask
show
jobs
submit
login
▲
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
(
terminal-bench-science.ai
)
23 points by
matt_d
2 hours ago
|
3 comments
add comment
akshay_akula 48 minutes ago
[-]
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
24 minutes ago
[-]
rubslopes 1 hours ago
[-]
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
vatsachak 43 minutes ago
[-]
Damn. These things aren't AGI... but I don't care.
Luna is good enough for me to give a parser spec and have it write one.
Rendered at 01:51:50 GMT+0000 (Coordinated Universal Time) with Vercel.
Luna is good enough for me to give a parser spec and have it write one.