🧔♂️ A friendly human may check it before it goes live. More news here
OpenAI launches new benchmark to test AI in freelance work
OpenAI has introduced SWE-Lancer, a new benchmark to assess large language models (LLMs) in freelance software engineering.
The dataset includes over 1,400 tasks from Upwork, with payouts exceeding US$1 million.
Tasks range from US$50 bug fixes to US$32,000 feature implementations and are categorized into independent engineering and managerial decision-making tasks.
Independent tasks undergo triple-verified evaluations by experienced engineers, while managerial tasks are assessed based on decisions from original engineering managers.
Testing of frontier LLMs against SWE-Lancer showed difficulty in solving most tasks.
To support further research, SWE-Lancer provides an open-source Docker image and a public evaluation subset, SWE-Lancer Diamond.
The initiative explores AI’s economic impact on software engineering by correlating model performance with financial outcomes.
The benchmark is now publicly available for researchers and developers.
Stay updated on the go with our mobile app.
Get latest insights with smoother, more personalized experience through TIA mobile app.




