Updates

Updates on our activities and progress.

Most recent
New models benchmarked: including GPT 5.5, Claude 5 Sonnet, Kimi K2.6, Gemini 3.5 Flash, DeepSeek V4 Pro, GLM 5.2, Qwen 3.7 Max, Grok 4.3, and Mistral Medium 3.5.
What's new in Phare V2 Phare V2 introduces a major update: a jailbreak module focused on circumventing safety guardrails to enable the generation of harmful content; and the inclusion of r...
We recently released the first large-scale evaluation using Phare, testing 17 leading language models across our three core safety modules: hallucination, bias & stereotypes, and harmful content generat...
Page of 1