ELITE STEM PhDs WANTED: FRONTIER AI EVALUATION
About the role
Next-gen AI models can solve standard textbook equations but fail at complex, multi-layered reasoning. Veron.ai is building frontier benchmarks to stress-test these boundaries. We need elite human minds to exploit these hidden AI logic gaps using complex, text-based problems that state-of-the-art models cannot crack.
Deliverables
The 3 Rules for your sample: Text-Only: Max 1,000 words. Use pure text or LaTeX equations. No diagrams, charts, or images. One Absolute Answer: The final solution must be a single number, string, or definitive value that a computer script can auto-grade. (No essays). The Trap: Include the step-by-step correct solution, and a 1-sentence note pointing out exactly where the AI’s logic will hallucinate and break down. Send your sample over, and if it passes our benchmark, we start contracting immediately