Abstract
A fundamental intent asymmetry plagues modern 3D asset creation: while state-of-the-art 3D toolchains demand precise, executable parameters, ordinary users typically provide vague, underspecified instructions. Current 3D agents treat this ambiguity as noise, defaulting to blind execution under a single-turn assumption. To address this limitation, we introduce CLARE, a clarification-aware and evolutionary 3D agent that treats intent asymmetry not as an execution error, but as an opportunity for strategic dialogue. By decoupling the generation pipeline into four specialized cognitive roles, CLARE intercepts and resolves underspecified instructions before invoking computationally expensive 3D tools to seamlessly execute tasks across five diverse domains: text-to-3D generation, single-view reconstruction, multi-view reconstruction, point cloud editing, and post-processing. Crucially, rather than relying on rigid manual rules, CLARE self-evolves its clarification policy via simulated multi-turn interactions. By optimizing a Multi-turn Reward, the agent internalizes the delicate balance between interaction efficiency and task completion. To rigorously test this, we construct 3D-Clarify, a comprehensive benchmark comprising 620 interaction scenarios with systematically injected ambiguity, missing information, and mistaken details. CLARE achieves state-of-the-art performance, with 60.40% and 43.34% success rates on single-step and multi-step tasks, respectively, more than doubling existing baselines. Both quantitative and qualitative results demonstrate that proactive clarification is the missing key to robust 3D execution.
Interactive trajectory
Clarify Before Executing
Replay how CLARE resolves an underspecified request, opens the execution gate, and produces a verified 3D tool call.
Task 3 / Single-step
Clarification dialogue
Method
CLARE Framework
CLARE separates decision gating, context refinement, executable synthesis, and feedback alignment, then improves its clarification policy through simulated multi-turn self-evolution.
Benchmark
3D-Clarify
A controlled benchmark for testing whether 3D agents can uncover hidden intent, request missing parameters, and correct mistaken constraints.
Experiments
Main Results
Goal Completion Rate (CR, %) and Task Success Rate (SR, %) on the 3D-Clarify benchmark. Best and second-best results are shown in bold and underlined, respectively.
| Method | 500 Single-step Tasks | 120 Multi-step Tasks | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ambiguity | Missing | Mistaken | Average | Ambiguity | Missing | Mistaken | Average | |||||||||
| CR | SR | CR | SR | CR | SR | CR | SR | CR | SR | CR | SR | CR | SR | CR | SR | |
| Zero-shot | 31.90 | 4.00 | 17.25 | 0.00 | 77.55 | 34.80 | 42.23 | 12.93 | 20.43 | 0.00 | 21.62 | 0.00 | 85.31 | 35.83 | 42.45 | 11.94 |
| Few-shot | 34.92 | 6.00 | 18.42 | 0.00 | 77.25 | 32.60 | 43.53 | 12.87 | 23.03 | 0.00 | 20.86 | 0.00 | 89.04 | 41.67 | 44.31 | 13.89 |
| CoT | 31.38 | 4.40 | 18.15 | 0.00 | 73.03 | 25.80 | 40.85 | 10.07 | 19.92 | 0.00 | 20.88 | 0.00 | 86.17 | 34.17 | 42.32 | 11.39 |
| ReAct | 29.88 | 5.40 | 22.42 | 1.60 | 50.38 | 16.40 | 34.23 | 7.80 | 14.51 | 0.00 | 18.04 | 0.00 | 74.72 | 13.33 | 35.76 | 4.44 |
| Reflexion | 39.47 | 8.40 | 28.07 | 2.40 | 69.31 | 37.60 | 45.62 | 16.13 | 22.53 | 0.00 | 25.60 | 0.00 | 83.68 | 23.33 | 43.94 | 7.78 |
| CLAMBER | 50.87 | 26.80 | 44.85 | 13.80 | 71.32 | 32.20 | 55.68 | 24.27 | 13.12 | 2.50 | 21.23 | 4.17 | 77.12 | 21.67 | 37.16 | 9.44 |
| CEP | 57.94 | 37.60 | 41.48 | 19.20 | 36.52 | 16.40 | 45.31 | 24.40 | 25.20 | 5.00 | 21.87 | 5.83 | 48.70 | 8.33 | 31.92 | 6.39 |
| 3D-GPT | 35.18 | 10.60 | 34.41 | 10.60 | 56.04 | 25.20 | 41.88 | 15.47 | 11.18 | 2.50 | 21.24 | 1.67 | 68.85 | 20.00 | 33.76 | 8.06 |
| CLARE-base | 59.31 | 45.80 | 57.58 | 41.00 | 81.90 | 59.00 | 66.26 | 48.60 | 38.16 | 23.33 | 40.73 | 22.50 | 81.25 | 35.83 | 53.38 | 27.22 |
| CLARE-SE-SFT | 62.29 | 49.00 | 69.47 | 52.00 | 85.43 | 66.80 | 72.40 | 55.93 | 49.72 | 34.17 | 68.10 | 44.17 | 83.77 | 48.33 | 67.20 | 42.22 |
| CLARE-SE-DPO | 70.52 | 60.00 | 75.32 | 59.40 | 74.68 | 61.80 | 73.51 | 60.40 | 53.41 | 34.17 | 68.74 | 49.17 | 72.75 | 46.67 | 64.97 | 43.34 |
CLARE-SE-DPO reaches 60.40% single-step SR and 43.34% multi-step SR, more than twice the strongest non-CLARE baselines.