← Home

About the project

This site supports exploration and verification of English–Chinese foul-word pairings derived from a parallel corpus. Data is enriched with keyword detection, optionally refined by a language model, and then verified by research assistants.

Data sources

The underlying corpus is built from bilingual sentence pairs (e.g. from OpenSubtitles or similar sources). Foul-word lists are applied to tag English and Chinese keywords; confirmed pairs are exposed here for search and visualization.

Methodology

The pipeline: (1) enrich raw corpus with foul-word regex matching, (2) optionally run an LLM to fill both EN and ZH keywords per pair, (3) import foul-match rows into the database, (4) research assistants verify or correct pairings, (5) confirmed pairs drive public search and visualizations.

The Team

This project is hosted by the Department of Translation at The Chinese University of Hong Kong.

This project is sponsored by Chung Chi College, The Chinese University of Hong Kong. (SAF Ref. No.: 2627A-046.)

Loading team…

References

For data and tooling references, see the project repository and documentation.