Running Reproduction: Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals 🎯 Explore experiment logs and collaborate with an AI agent