Developer Toolsjetbrains_blogAlphaLab AI score 26/100
LLMs Show Divergent Paths Despite Equal Task Resolution
In a benchmark test, Claude Opus 4.7 and Gemini 3.5 Flash solved the same number of coding tasks, but their execution profiles revealed significant differences. Opus averaged 184 steps and cost $2.79 per run, while Gemini took 271 steps but cost only $1.24. This underscores the limitations of resolve rate as a metric, which only measures task completion without detailing the process. JetBrains' Junie coding agent evaluates models based on functional outcome, execution efficiency, patch quality, and process quality, offering a more comprehensive assessment. The findings highlight the need for nuanced evaluation methods that consider both result and trajectory.