input or output path. That lets you apply moderation before model invocation, after generation, or both.
What moderation covers
Typical moderation coverage includes hate and harassment, sexual content, violence and self-harm, illicit guidance, and other categories your policy treats as disallowed for the product.Common moderation categories
Moderation categories are never purely theoretical. They need to reflect your product, geography, and risk posture. A customer-support assistant, an internal operations assistant, and a moderation pipeline for user-generated content will not all draw the same line in the same place.
Where it helps most
This guardrail matters on both sides of the model. On theinput path, it limits what the model is asked to process. On the output path, it limits what the application is willing to return or automate. It is especially important when your application rewrites, summarizes, translates, or otherwise transforms user-supplied material.