Concrete job: Grok reads the tweet and stamps a label (reply-spam, adult, graphic, self-harm, terms of service…).
You don’t see Grox. You see the consequence later: Visibility Filtering uses that stamp to show, blur, or hide.
Twelve plan names are public. The actual prompts are not — we know what it is asked to judge, not the wording.
One published number: a reply-spam score of 0.97 or more becomes a “risky high-viz reply.” That is how a pasted or spammy reply can die for strangers on For You even if followers still see it.
Grox does not rank. It only labels.
More: Show, blur, or hide.