NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows (terminal-bench-science.ai)
akshay_akula 48 minutes ago [-]
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
24 minutes ago [-]
rubslopes 1 hours ago [-]
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
vatsachak 43 minutes ago [-]
Damn. These things aren't AGI... but I don't care.

Luna is good enough for me to give a parser spec and have it write one.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 01:51:50 GMT+0000 (Coordinated Universal Time) with Vercel.