Google's Gemini 3.6 Flash scored 56.1% on the updated WeirdML v2 benchmark, performing worse than its predecessor 3.5 Flash and significantly behind frontier models. The model frequently times out (50% of the time, up from 33% for 3.5 Flash) due to miscalibrated execution time estimates and an inability to adjust its plans based on direct feedback. The WeirdML v2 update includes 19 tasks, API cost tracking, and metadata, showing a varied Pareto frontier with 11 models from six companies achieving optimal accuracy for given costs.
Google's latest Flash model struggles with basic self-correction, indicating a fundamental limitation in its agentic planning and execution capabilities.