A text moderator that classifies prompts and responses using a defined hazard taxonomy; this profile examines Llama Guard 3–8B.
Use Llama Guard to ask whether a conversation contains content covered by a moderation policy. It is a separate model in the application’s checking pipeline. Deciding whether a flagged request should be blocked, redirected or reviewed remains an application decision.
Meta describes an 8B model that generates safe/unsafe classifications and violated categories for prompts or responses. The documented taxonomy has 14 categories, including code-interpreter abuse.
Source: Meta · Llama Guard 3–8B model cardThe card lists English, French, German, Hindi, Italian, Portuguese, Spanish and Thai. Language coverage does not establish equal accuracy across dialects, mixed-language conversations or specialist terminology.
Source: Meta · Llama Guard 3–8B model cardOur recommendation: retain the category and relevant context for review, distinguish input from response classification, and specify what happens if the moderator times out or returns an unparseable result.
In a support assistant, screen the proposed answer before showing it. If it is flagged, offer a safe alternative or human review. Include legitimate troubleshooting language in testing so ordinary technical requests are not mistaken for abuse.
A content label does not establish factual correctness, grant tool authority or guarantee a private-data boundary. Do not transfer this text checkpoint’s specifications to other Llama Guard releases.
How to read an AI safety evaluationThe sources behind this page, with a reason to open each one. Practical examples and recommendations are our editorial interpretation.
Defines the moderation labels, supported languages, evaluation setup and limitations of this checkpoint.
Pairs safe prompts with unsafe contrasts to investigate unnecessary refusals. Historical model results are not current rankings.
Sources reviewed 13 September 2026. Product documentation can change. How we use evidence