📰 Latest Update: Benchmarking 17 LLMs with Phare (June 2025)
Published on June 16, 2025 by Stanislas Renondin

 
We recently released the first large-scale evaluation using Phare, testing 17 leading language models across our three core safety modules: hallucination, bias & stereotypes, and harmful content generation. 
 
🧠 Key Findings
 
  • Hallucination remains a major weakness: Many models are sensitive to how prompts are framed. When users present false claims with high confidence, models are more likely to agree, a phenomenon known as sycophancy. Additionally, when instructed to be brief, several models show degraded performance on factuality and misinformation tasks.
  • Biases are persistent and subtle: Even without explicit prompts, models reproduce harmful stereotypes. For instance, manual labor is systematically associated with male characters, and ethnic/religious attributes correlate with social roles. Most models recognize these associations as biased when explicitly asked — but still reproduce them in open generation.
  • Harmful content is better mitigated: Across the board, models scored between 70% and 100% resistance to harmful prompts related to eating disorders, substance misuse, or dangerous advice. Newer models like GPT‑4o and Claude 3.7 Sonnet performed best, showing the progress made on this front.

📊 Technical Summary

 
  • Modules evaluated: Hallucination (factuality, misinformation, tool use), Bias & Fairness (attribute co-occurrence and self-coherency), Harmfulness (vulnerable framing, sycophancy).
  • Languages covered: English, French, Spanish.
  • Evaluation method: LLM-as-a-judge (3-model voting), with human validation and statistical significance checks (chi-squared, Cramér’s V).

📂 Resources


Â