safetydesk.io

What is allowed

Fiction, without restriction.

Safety Desk is the moderation operation behind an adult generative-image service. Its starting position is unusual and deliberate: explicit sexual content between consenting adults is lawful, it is what people come for, and it is not something to be quietly throttled, blurred or refused on a hunch.

Most moderation systems for this material work by narrowing what can be asked for until the product barely functions, then calling the wreckage safety. This one does the opposite. It says yes to fiction — every fantasy, every scenario between adults, every style — and spends its entire effort on a short list of things that are never fiction, never negotiable, and never shipped.

What is refused, always

Twelve limits, and no discretion about them.

These are absolute. No setting relaxes them, no volume of custom justifies them, and no commercial argument reaches them.

Every generation passes three screens before anyone sees it

They run in order, and they are independent on purpose: a failure in one is caught by the next, and the first of them does not depend on a model, a network or a vendor being available.

Screen one

A deterministic floor

A fixed list of prohibited terms is applied to the prompt in our own code, offline, before any provider is contacted and before any credit is spent. It has no model behind it and cannot be talked around, degraded by an outage or improved by a jailbreak. It is the only part of the system that is not probabilistic.

Outcome: refused outright.

Screen two

The prompt is classified

A vision-language classifier reads the request against a written policy of twenty-four categories and returns its categories, a confidence and its own reasoning in about a second. A category only acts when the classifier substantiates it — naming the person, naming the animal, naming the basis for non-consent. An unevidenced label is recorded and ignored.

Outcome: refused for the absolute categories, otherwise held for a person.

Screen three

The picture itself is classified

Words are not the artifact. The finished image — or the clip, sampled a frame a second — is classified before it is delivered, so what is judged is what was actually produced rather than what was asked for. Until that screen returns, the customer is still looking at a progress state.

Outcome: withdrawn, or held, or delivered.

A machine may hold something. Only a person may release it.

A classifier flag never becomes an automatic refusal. When something is flagged, the result is withheld behind a plain notice telling the customer a person is reviewing it and roughly how long that takes. It is not deleted, not refused, and not silently downgraded to something blander.

A reviewer then approves it — in which case it appears immediately — or rejects it, citing which of the twelve clauses it conflicts with. That clause is what the customer is told, in the same words the policy uses. Nobody receives a vague refusal, and no refusal cites a rule that does not exist.

Reviewers work in a separate console on a separate domain, behind their own sign-in, with no access to the product. Work is claimed before it is decided so two people cannot rule on the same item, and every decision carries the name of the person who made it.

Access is graded. A contractor sees the ordinary review queue. Refusals and reinstatements need seniority. Anything touching the minor category is a separate, separately cleared queue that most staff can never open — they can see that such an item exists on an account, which is what they need to judge intent, and never its contents.

Suspending a person is a different power from deciding one picture, and reviewers do not have it. It requires seniority, a written reason every time, and it is recorded against the name of whoever did it.

When the classifier cannot be reached, the system allows and counts it rather than refusing. An outage should degrade screening, not confiscate work people paid for — and the deterministic floor is still standing. Those events are counted separately from screened traffic and are never presented as coverage.

Every control, and whether it is on

The full set, and the state each one is in. Nothing here is aspirational: if a control is listed, it is running on every request today.
ControlStateWhat it does
Prohibited-terms floor Always on Offline, deterministic, unaffected by outages. Cannot be switched off.
Prompt classification On Twenty-four categories, evidence-gated, roughly one second per request.
Output classification On Images and clips, before delivery. Clips are sampled a frame a second.
Hold instead of refuse On A flag withholds the result pending human review. It never refuses on its own.
Clause-cited notices On Every refusal, hold and removal names the clause it rests on, in the customer's own copy.
Stated review time On Shown on every held result. Set by a lead; two minutes by default.
Graded queue access On Four roles. Minor-class material is a separately cleared queue.
Decision ledger On Every screening and every human decision, with categories, confidence, reasoning, latency and reviewer. Exportable.
Account suspension On Senior and lead only, written reason required, optionally withdrawing everything the account has produced.
Curated example prompts On, bounded Nine fixed example prompts written and vetted in-house are recorded but not acted on, so a one-tap example is never held. The floor and the output screen still apply to them, and one edited word removes the exemption.

The twelve clauses, in the words customers are given

When something is refused, held or removed, one of these sentences is what the person sees. The console composes the notice from this same list, so the policy and the customer's copy cannot drift apart.
AMinorsContent that depicts or implies anyone under 18 is prohibited without exception.
BNon-consentContent depicting non-consensual sexual activity — including anyone asleep, drugged, unconscious or coerced — is prohibited.
CReal peopleSexual or nude depictions of real, identifiable people, including lookalikes and deepfakes, are prohibited.
DReal family abuseSexual content involving real family members, or family-framed content with a minor or a person who cannot consent, is prohibited.
EBestialitySexual activity between a person and a real animal is prohibited.
FDeath and extreme harmSnuff, necrophilia, mutilation, torture, heavy blood or gore, and glorified self-harm or suicide are prohibited.
GTraffickingContent promoting, advertising or depicting sex trafficking or exploitation is prohibited.
HWeapon coercion and terrorismWeapons used to threaten, force or harm a person in a sexual context, and sexualised terrorism or mass violence, are prohibited.
IBodily wasteScat, urine or vomit on a person presented sexually is prohibited.
JHate and harassmentSexual content that dehumanises or incites violence against people for who they are is prohibited.
KDrugsPromotion, facilitation or depiction of illegal drug use, sale or manufacture is prohibited.
LIntellectual propertyUse of trademarks, logos, counterfeit goods or a protected character or likeness for commercial gain is prohibited.

For banks, acquirers and payment partners

The point of building it this way is that nothing here has to be taken on trust. Every screening decision and every human decision is a row in a ledger with the categories, the classifier's own reasoning, its confidence, the latency, the outcome and, where a person was involved, their name and the clause they cited.

On request we can provide the written classifier policy verbatim, the category lists and what each one does, coverage over any period with fail-open events counted separately, an export of human decisions with clauses and reviewers, and the complete evidence for any individual case.

Reviewers and named partner contacts work in the console at console.safetydesk.io, behind organisational sign-in. It carries the same written policy this page describes, generated from the running system rather than maintained by hand — so what a reviewer enforces and what a partner is shown are the same document.

For access or an export, write to safety@generateporn.ai.