Characterizing Bluesky Content Moderation Service

From automation of service to the landscape of harms it detects

The first large-scale audit of Bluesky's default labeler, built from 10.6M real moderation labels streamed publicly over AT Protocol in 2025.

Abstract

Empirical research on content moderation is fundamentally constrained by the opaque deployment of moderation systems on major social media platforms. To this end, the recent emergence of decentralized platforms with transparent, public moderation logs presents an unprecedented opportunity for independent audits. In this work, we leverage this architectural transparency to conduct the first large-scale audit of the default moderation system on Bluesky, the Bluesky Moderation Service (BMS). Analyzing its 10.6M moderation labels from 2025, we investigate three foundational aspects: (i) its mechanism (the degree of automation versus human oversight), (ii) its efficacy (accuracy in detecting harms), and (iii) its purpose (the landscape of harms it identifies).

Our findings reveal a human-AI collaborative system where labels for sexual and graphic content are applied automatically in seconds, while nuanced and high-stakes labels require more human oversight, taking hours or days. Through a manual annotation study, we find the BMS operates with high precision (0.837), but struggles with low recall (0.222), with our annotators identifying 4.5× more harmful content than the moderation system in a random sample. Finally, unsupervised clustering of the most frequently applied labeled posts uncovers detected harms ranging from hostility in discourse toward protected groups to the spread of sexually explicit and other graphic content.

RQ1 · Mechanism

How much of BMS is automated?

We use labeling delay — the time between a post's creation and its label — as an empirical proxy for automation: labels applied within seconds, at scale, are likely automated; labels taking hours or days indicate human review. This is computed directly from BMS's 2025 label stream (10.46M valid records).

Median labeling delay vs. number of posts labeled

Each point is one label. Both axes log-scaled. Hover for exact percentiles. Color = automated (teal) vs. human-oversight (rose), split by the delay dichotomy we observe.

Automated labels: CDF of delay by embed type

porn, sexual, nudity, self-harm, graphic-media, spam, !warn, !hide.

Human-oversight labels: CDF of delay by embed type

intolerant, rude, threat, sexual-figurative, !takedown.

Redressal latency: delay until a label is negated/reversed

Median time between post creation and label removal, for the 206,724 negated labels we observed.

RQ2 · Efficacy

Precision and recall of the BMS

Annotators independently labeled 1,000 BMS-flagged posts (labeled set, for precision) and 1,000 random firehose posts (random set, for recall).

Unsafe content: labeled set vs. random firehose set

Precision 0.837 — of posts BMS flags, 83.7% are genuinely unsafe by human judgment. In the random firehose sample, every post BMS flagged (6/6) was confirmed unsafe (100% precision there too).

Recall 0.222 — of the 27 unsafe posts human annotators found in a random 1,000-post sample, BMS caught only 6. Automated-label recall is 0.600; recall for manually-applied labels in this sample is 0 — moderators only review reported/surfaced content.

→ Annotators found 4.5× more harmful content than BMS flagged, in a random, unbiased sample of the firehose.

RQ2, continued · Blindspots

Why does the automated pipeline miss content?

BMS's automated pipeline calls Hive AI (128 classes across 54 heads) for every image, then Automod — Bluesky's open-source rule engine — applies hard-coded thresholds to just 16 of those classes. We matched 336 BMS-labeled posts to their nearest unlabeled neighbor in a 40M-post firehose sample (Qwen3-VL-Embedding-2B + faiss) to probe what the pipeline misses.

Step 1

Hive AI

Returns 128 class-score pairs (54 heads) per image, e.g. yes_sexual_activity: 0.91.

Step 2

Automod rules

Hard-coded thresholds on 16 of 128 classes, e.g. yes_self_harm ≥ 0.96self-harm.

Step 3

BMS label

Priority cascade for sexual content: pornsexualnudity.

H1: close threshold misses

Hive scores on Automod-mapped heads for the 22 missed posts, sorted by median gap to threshold (red marker). Real per-post scores.

H2: rule-set gap

Hive scores on heads Automod never reads, for the 14 missed posts. Automod takes no action despite near-perfect scores.

RQ3 · Purpose

The landscape of harms

For each label, posts are embedded with Qwen3-VL-Embedding-2B (text + media jointly), projected to 5D with UMAP, and clustered with HDBSCAN (hyperparameters tuned per label via Optuna/TPE, selected by DBCV score). Clusters are named by a vision-language model (Qwen3-VL-32B-Instruct) given sampled posts, validated by three human annotators (92.5% rated appropriate/very appropriate, Krippendorff's α = 0.686). Points below are the first two UMAP components for a downsampled set of real posts per cluster.

    Clusters within a label

    Click a cluster — in the chart or the legend above — for details.

     

    Harms to protected groups & civic discourse

     intolerant, rude, threat surface gender/LGBTQ+, religious, racial and political targeting — from slurs to explicit death wishes directed at political figures.

    Harms to physical & mental well-being

    graphic-media and self-harm capture war/atrocity imagery, historical execution photos repurposed politically, and cutting/self-injury content, often under hashtags like #sh.

    Harms to minors from sexual content

    Four labels — porn, sexual, nudity, sexual-figurative — span a spectrum from commercial adult content to fine-art nudity to stylized/anime erotica.

    Reproducibility

    Data & code

    Code and data are available at github.com/iampushpdeep/BMS.

    Team

    Pushpdeep Singh

    MPI-SWS
    Germany

    Website

    Sayeh Jarollahi

    Saarland University
    Germany

    Website

    Ayan Majumdar

    MPI-SWS
    Germany

    Website

    Vabuk Pahari

    MPI-SWS
    Germany

    Website

    Abhijnan Chakraborty

    IIT Kharagpur
    India

    Website

    Abhisek Dash

    MPI-SWS
    Germany

    Website

    Krishna P. Gummadi

    MPI-SWS
    Germany

    Website

    Ingmar Weber

    Saarland University
    Germany

    Website

    Contact

    Pushpdeep Singh

    psingh@mpi-sws.org

    BibTeX

    @misc{singh2026characterizingblueskycontentmoderation,
          title={Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms}, 
          author={Pushpdeep Singh and Sayeh Jarollahi and Ayan Majumdar and Vabuk Pahari and Abhijnan Chakraborty and Krishna P. Gummadi and Ingmar Weber and Abhisek Dash},
          year={2026},
          eprint={2609.11373},
          archivePrefix={arXiv},
          primaryClass={cs.CY},
          url={https://arxiv.org/abs/2609.11373}, 
    }