Project 6.21: Monitoring Administrative Persecution Targeting LGBTQIA+ Expression in Russia

We monitor persecution targeting LGBTQIA+ expression in Russia. This dataset results from our monitoring. We collected incidents of administrative persecution against individuals and organisations for the so-called “LGBT propaganda” from open sources: court documents, mass media and social media. Having such incidents in a structured form enables us to investigate connections between cases and identify patterns of persecution. Others can freely reuse this dataset. Read more in the detailed documentation below.

Detailed documentation

Project 6.21 Dataset Documentation

1. General information

  1. Title: Project 6.21: Monitoring Administrative Persecution Targeting LGBTQIA+ Expression in Russia
  2. Creator: Grey Rainbow project, Dataout Foundation (Dataout)
  3. Persistent Identifier: 10.5281/zenodo.22693416
  4. Published on: Zenodo
  5. Publication Date: 10.09.2026
  6. Version: 1.0.0
  7. Last Updated: 10.09.2026
  8. Associated File Formats: JSON-LD, RDF/XML, N-Triples, Turtle (TTL), CSV
  9. Licence: Creative Commons Attribution 4.0 International
  10. Contact: greyrainbow@dataout.org
  11. This documentation was updated on: 10.09.2026
  12. Cite as: Dataout Foundation. (2026). Project 6.21: Monitoring Administrative Persecution Targeting LGBTQIA+ Expression in Russia (Version 1.0.0) [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.22693416

2. Motivation

  1. The Grey Rainbow project monitors incidents of persecution targeting LGBTQIA+ expression in Russia, in which individuals and companies are punished by courts for the so-called “propaganda of non-traditional sexual relations and/or preferences, sex change, or refusal to have children” or for distributing such information among minors. This dataset aggregates cases collected during our monitoring. We use this dataset for two purposes: (1) to preserve monitored cases together with their provenance information and related sources in a structured form; and (2) to investigate connections between cases and identify patterns of persecution. The dataset can be freely reused by researchers, journalists, human-rights organisations, and others. Our goal is to make evidence of human rights violations more traceable and reusable.

  2. This dataset was created by the Grey Rainbow project at Dataout Foundation (Dataout).

  3. Dataout did not receive funding to carry out work on this dataset. This dataset is the result of volunteer work.

3. Composition

  1. This dataset results from the monitoring activity carried out between 01.12.2025 and 29.05.2026. We collected 316 monitoring records covering incidents from 2014 to 2025. A single monitoring record consists of provenance information and content. The provenance information includes the following:

    • Sources from which the record was derived. There are several source categories. The primary sources of the monitoring records in this dataset are court documents. Other source categories include mass-media publications, social-media posts by court representatives, and press releases by governmental bodies. A single record may have multiple sources. Sources in the 'court document' category contain metadata such as their category, creator, document publication date, document number (assigned by a court), and unique ID (assigned by a court). They may also contain the URL and an archived URL of the page on a court website where the document was published. The creator of a court document is another instance: a court. A court has a court code (used on its official website), a name, and a website URL. Sources belonging to other categories contain their category, creator, URL, and archived URL.
    • Monitoring record metadata, including its number, identifier, and the monitoring activity from which it originates (an instance). A monitoring activity contains a name, description, category (in this case, 'Administrative Persecution'), start date, end date, a connection to an organisation (an instance) that carried out the activity, and a list of all monitoring records generated by the activity.

    The record's content consists of information extracted from sources as well as external information:

    • Region (an instance) in which administrative proceedings were initiated, represented by a region code, name, abbreviation, ID, and mapping to GeoNames.
    • Year (a date-typed string) in which a monitored incident occurred, that is, when a court case was heard.
    • Defendant (an instance) against whom the administrative proceedings were initiated, containing a category (individual, official, or legal entity) and an ID. Defendants in the 'legal entity' category also contain a legal name. Defendants in the 'official' category have names and may contain an affiliation, indicating that an official is associated with an organisation that may act as another defendant.
    • Administrative article (an instance) imposed by the court. In this dataset, there are two articles: Article 6.21, which includes eight parts, and Article 6.21.2, which includes four parts. We link each record to an article part. Each part has a name and description in English and Russian.
    • Trigger object (an instance) representing a piece of information or an action that triggered the persecution. Each trigger object has an ID and may also have a category, name, description, temporal information (year), and URL.
    • Penalty (an instance) imposed by a court, with optional category and value. Penalties categorised as 'fine' have a value for the monetary amount and currency. Penalties categorised as 'detention' have a value for the number of days; penalties in other categories do not have values.
    • Snippet (a text string) summarising an incident in English and Russian; it may include quotations from sources.
    • External information, including URLs to mass-media or social-media publications that mentioned an incident and IDs of trigger objects (such as films or books) from external sources, for example, The Movie Database, YouTube, or Open Library.
  2. We provide this dataset as Linked Open Data. We developed the Dataout Ontology and represented the records in a Knowledge Graph (KG). The KG files are available on our GitHub in JSON-LD, RDF/XML, N-Triples, and Turtle. Refer to the Dataout Ontology documentation to learn about the connections between the instances. The ontology documentation also includes documentation of the controlled vocabularies we used to categorise instances. The KG contains more than 13,500 triples. To evaluate the KG and demonstrate its analytical value, we formulated SPARQL queries (see the Jupyter Notebook queries.ipynb). We described the processes of data collection, ontology development, and KG development in the paper "Using Linked Data for Monitoring Administrative Persecution Targeting LGBTQIA+ Expression".1

  3. We provide a simplified version of this dataset in a CSV file. It offers an overview of the dataset but does not contain all the information or relationships between instances represented in the KG. The CSV contains one row per monitoring record and the following 14 columns:

    • record_id: monitoring record ID (assigned by us);
    • case_number: a court case number as found in the document (assigned by a court), not globally unique;
    • year: year of the monitored incident;
    • region: code of the Russian region in which the incident occurred (where administrative proceedings were initiated);
    • court_code: court identifier (used by official court websites);
    • article: administrative article and its part associated with the case;
    • defendant_id: defendant ID (assigned by us);
    • defendant_category: category of the defendant, such as individual, official, or legal entity;
    • penalty: penalty category and, where applicable, its value (given in brackets); there may be multiple penalties separated by semicolon;
    • trigger_object: a combination of a trigger object category and its ID (given in brackets); there may be multiple trigger objects separated by semicolon;
    • source: a combination of a source category and its URL or unique ID (given in brackets); there may be multiple sources separated by semicolon;
    • mentioned_in: URL or identifier of an additional source mentioning the case, where available; there may be multiple URLs separated by semicolon;
    • snippet_en: English-language summary of the case (written by us);
    • snippet_ru: Russian-language summary of the case (written by us).

An example of a record from the CSV file:

record_id case_number year region court_code article defendant_id defendant_category penalty trigger_object source mentioned_in snippet_en snippet_ru
record_6679c2c9 5-28/2014 2014 27 27MS0026 6.21_2 official_0002 official fine (50000) Mass media (mass_media_001) court https://web.archive.org/web/20220304091520/https://www.bbc.com/russian/rolling_news/2014/01/140131_rn_khabarovsk_gay_scandal Snippet text EN Snippet text RU
  1. This dataset is not complete, meaning that it does not contain all incidents of administrative persecution. This dataset includes only incidents that resulted in courts issuing punishments and for which we could obtain sufficient evidence. For example, we did not include cases that were dismissed by courts or for which information was unavailable on the official court websites.

  2. There is missing information in this dataset. Not all incidents are described with complete information. Each monitoring record contains at least the following information: case number, year, region, administrative article, defendant, penalty, trigger object, source, and a textual snippet in English. Other information may be missing. There are three reasons for missing information:

      1. The information was not available in the sources. We indicate which information was unavailable in the textual snippet for each record. This missing information is usually related to penalties and trigger objects because courts redacted such information. We explicitly created instances (URIs) of penalties and trigger objects even when they were unclear or unknown. A penalty may not have a category or value (for example, the amount of a fine) if this information is unknown. A trigger object may not have a category if its category is unknown. We extracted trigger objects based on information in the sources for each incident when they were clearly mentioned. For example, if a court decision states that an individual was fined for publishing three images on social media, we created three trigger-object instances with the category “Social Media”; when the number or nature of the trigger objects was unclear, we created one trigger-object instance.
      1. Some information is not applicable to particular instances. Instances belonging to different categories may be described with different information. A trigger object may not have a name, description, temporal information, or URL because these details are not always applicable to every trigger-object category. Penalties may not have values because they belong to categories without values, such as 'warning' or 'expulsion'. Sources in the 'court' category (court documents) may not have a name, URL, or archived URL. Sources in other categories are not connected to a court ID, court document number, unique court document ID, or court document because these properties are not applicable.
      1. The information is incomplete and may be added in future releases of this dataset. Missing information includes Russian-language record snippets (this version includes 73 RU snippets), archived URLs, and connections to external publications that mention incidents (their URLs and/or archived URLs).
  3. This dataset is self-contained, meaning that it does not rely on external sources in order to be processed. It contains external links that enrich the information about records, including links to original sources, publications that mention incidents, and sources about trigger objects. Where possible, we provided archived URLs using the Internet Archive’s Wayback Machine. We also use reference datasets about Russian courts, regions, and administrative articles. These datasets are included in the KG and are also provided in a separate GitHub repository.

  4. This dataset does not contain confidential information.

  5. This dataset contains offensive and stigmatising words and phrases. These occur particularly in snippets and quotations from court documents. The quotations are retained to illustrate the language and position of authorities; they represent neither the views of Dataout nor those of the creators of the dataset.

  6. This dataset relates to people. It contains information about individuals who appear as defendants in administrative proceedings. We did not include information that can be used to identify individuals (such as social-media handles or nicknames), but in some cases, we preserved the names of individuals who act as company representatives or are public figures, such as activists or journalists.

4. Data collection process

  1. We collect data on incidents of administrative persecution targeting LGBTQIA+ expression in Russia from open sources. Information about each incident was extracted from textual sources, such as court documents, mass-media and social-media publications, and websites. In addition to the extracted information, we wrote summaries of cases (snippets) in English and Russian.

  2. In this version of the dataset, we collected incidents in which courts issued punishments based on Articles 6.21 and 6.21.2 of the Code of Administrative Offences of the Russian Federation.

  3. The primary sources are official court websites. We used several strategies to collect data on administrative persecution. First, we retrieved all available court cases concerning Articles 6.21 and 6.21.2 from the official websites using custom-developed software; on some websites that could not be parsed automatically, we used their internal search functionality. Second, we consulted related works to identify previously known cases of persecution.23 Third, we used keyword searches on search engines to find mass-media and social-media publications mentioning cases of persecution. When we found a case in a source other than an official court website, we verified whether the case had been registered by a court. Sometimes, courts publish information about administrative proceedings without publishing the associated full-text documents (decisions). We used the paid legal database "Consultant Plus"4 and a closed database operated by OVD-Info5 to obtain such documents.

  4. To record each incident, we read all the associated documents and publications we identified, extracted information from them, registered the provenance information, and wrote a summary. We extracted information from textual sources manually. The metadata of court documents published on the official websites (for example, their publication dates and unique IDs) was extracted automatically. We collected all monitoring records, including their provenance and content, in a tabular format. We then automatically assigned unique IDs (URIs) to records, defendants, penalties, trigger objects, and sources. The resulting tabular dataset was converted into the KG and the simplified CSV. The conversion process is in the notebook generate_kg.ipynb.

  5. Before collecting the data, we established potential risks and ethical principles. We do not publish the full texts of court documents, which may contain personal information; instead, we provide anonymised summaries of cases. We state explicitly that the collected incidents represent state persecution targeting LGBTQIA+ expression. We retain quotations from court documents in case summaries that contain stigmatising and offensive terms and phrases to illustrate the position of authorities and their legal language.

5. Uses

  1. We use this dataset for the purposes outlined in Section 2, paragraph 1. All our work related to this dataset will be published on the Grey Rainbow project website: greyrainbow.dataout.org.

  2. Others can share and adapt this dataset under the CC BY 4.0 licence, provided that they give appropriate credit and indicate whether changes were made.

  3. When reusing this dataset, its limitations (Section 3, paragraph 4) and sensitivities (Section 3, paragraphs 7 and 8) should be considered and acknowledged.

6. Maintenance

  1. This dataset is published on Zenodo, GitHub, and the Dataout website; it is supported and maintained by the Grey Rainbow project at Dataout.

  2. We plan to update, extend, and modify this dataset and its documentation. Each version of the dataset will have a unique identifier and version number.

  3. The ontology and controlled vocabularies, on which the KG relies, are maintained separately and are documented at dataout.org/ontology.

  4. For questions and suggestions, contact greyrainbow@dataout.org. To report errors, open an issue on GitHub. To contribute related data or derivatives based on this dataset, upload a submission to the Dataout Zenodo Community.

References

When compiling this documentation, we reused the framework suggested in "Datasheets for datasets" by Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM 64, 12 (December 2021), 86–92. https://doi.org/10.1145/3458723


  1. Using Linked Data for Monitoring Administrative Persecution Targeting LGBTQIA+ Expression. Nesterov, A., Katsuba, S. (2026) Posters, Demos, Blue Sky, and Tutorials at SEMANTiCS 2026, Sep 2026, Ghent, Belgium. To be published. 

  2. Citizens’ Watch and Sphere Foundation, Judicial harassment against LGBT+ under ‘propaganda’ law in Russia, Online, 2024. URL: https://spherequeer.org/judicial-harassment/

  3. S. Katsuba, Russian “Gay Propaganda Law”: A Comprehensive Qualitative Analysis of the Legislation and Case Law, Problems of Post-Communism 73 (2026) 138–151. doi:10.1080/10758216.2025.2487577

  4. https://www.consultant.ru 

  5. https://ovd.info/en