Blog

Updates and announcements from the LLMEval team.

·LLMEval Team

LLMEval-Logic Is Now Open Source: Code, Data, and Leaderboard

LLMEval-Logic is a Chinese logical reasoning benchmark double-audited by the Z3 SMT solver and human rubrics, and toughened via an adversarial-hardening agent loop. The code, complete evaluation pipeline, and 80% of the data are now open source (196 Base + 154 Hard + 196 rubrics); the remaining 20% is held out as a private contamination-resistant test set maintained by Fudan NLP Lab.

LLMEval-Logicopen sourcelogical reasoningZ3contamination-resistant
Read more