Last gate: show, blur, or gone. A hide cannot be ranked back in.
Who writes the labels
Other systems mark the post first. Visibility Filtering only reads the marks.
- Grox — Grok-based plans: terms of service, adult, graphic, self-harm, reply-spam, and more. The plan names are public. The prompts are not.
- Scarecrow — simple rules on the tweet. Twenty files are public: copy-pasta, bad URLs, NSFW, gore, and a few model scores.
- Agatha — models on the tweet and the account. High scores become NSFW, spam, or abuse labels.
- AES — adult scores on images and video.
- Reputation — used for a “high rank” exemption on some spam rules.
What happens then
Everyone shares a base set of rules. Strangers on For You get 27 extra drops: high-recall spam, high-recall NSFW, abusive, NSFW avatar, and more. Followers can still see those posts. Recommendations to strangers cannot.
A Grox reply-spam score of 0.97 or more becomes a “risky high-viz reply.” That threshold is in the public rules.
Some severity cutoffs are blanked out in the files. Treat them as unknown. The blur look is not in this repo either — only the three-way decision.
<aside>
✅
What helps
- Clean media and a clean profile. Fewer adult and NSFW-avatar labels for strangers.
- A real conversation, not pasted replies. Copy-pasta and reply-spam are the published traps.
- High reputation. Cred 54+ (or 25k+ followers) skips some spam rules — never the worst safety ones.
</aside>
<aside>
🚫
What hurts
- NSFW or gore. Dropped for strangers on For You.
- Copy-pasta and reply spam. Especially Grox ≥ 0.97.
- Bad or unsafe links. Public URL rules mark those.
- Hate, violent speech, abuse, civic integrity. Dropped even for followers. Reputation does not save this.
</aside>
<aside>
→
Next: Ads and extras
</aside>