1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
output_label = response["choices"][0]["text"] # This is the probability at which we evaluate that a "2" is likely real# vs. should be discarded as a false positivetoxic_threshold = -0.355if output_label == "2": # If the model returns "2", return its confidence in 2 or other output-labels logprobs = response["choices"][0]["logprobs"]["top_logprobs"][0] # If the model is not sufficiently confident in "2",# choose the most probable of "0" or "1"# Guaranteed to have a confidence for 2 since this was the selected token.if logprobs["2"] < toxic_threshold: logprob_0 = logprobs.get("0", None) logprob_1 = logprobs.get("1", None) # If both "0" and "1" have probabilities, set the output label# to whichever is most probableif logprob_0 isnotNoneand logprob_1 isnotNone: if logprob_0 >= logprob_1: output_label = "0"else: output_label = "1"# If only one of them is found, set output label to that oneelif logprob_0 isnotNone: output_label = "0"elif logprob_1 isnotNone: output_label = "1"# If neither "0" or "1" are available, stick with "2"# by leaving output_label unchanged.# if the most probable token is none of "0", "1", or "2"# this should be set as unsafeif output_label notin ["0", "1", "2"]: output_label = "2"return output_label
We generally recommend not returning to end-users any completions that the Content Filter has flagged with an output of 2. One approach here is to re-generate, from the initial prompt which led to the 2-completion, and hope that the next output will be safer. Another approach is to alert the end-user that you are unable to return this completion, and to steer them toward suggesting a different input.
Is there a cost associated with usage of the content filter?
No. The content filter is free to use.
How can I adjust the threshold for certainty?
You can adjust the threshold for the filter by only allowing filtration on the labels that have a certainty level (logprob) above a threshold that you can determine. This is not generally recommended, however.
If you would like an even more conservative implementation of the Content Filter, you may return as 2 anything with an of "2" above, rather than accepting it only with certain logprob values.output_label
How can you personalize the filter?
For now, we aren't supporting finetuning for individual projects. However, we're still looking for data to improve the filter and would be very appreciative if you sent us data that triggered the filter in an unexpected way.
What are some prompts I should expect lower performance on?