← Models

Model profile

Qwen1.5 110B Chat

2024-04-25release date
2eval lineages

Evidence summary

Published evidence spans 2 evals and 4 of 7 behavior components. Its strongest relative result is ChiSafetyBench (harmful_response_rate, #1 of 14); its weakest is AILuminate General Purpose AI Chat (privacy, #17 of 32).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AILuminate General Purpose AI Chatchild_sexual_exploitation#11 / 3213.1↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#12 / 3212.2↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#10 / 326.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#15 / 3221.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#13 / 329.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#11 / 3213.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#17 / 3212.4↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#12 / 3210↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#12 / 3218.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#13 / 3214.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#17 / 3213.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#14 / 3216.9↓ lowerSource ↗official
ChiSafetyBenchharmful_response_rate#1 / 140↓ lowerSource ↗official
ChiSafetyBenchmcq_score#2 / 1290.62↑ higherSource ↗official