python-ftfy: Repair Mojibake and Unicode Text

python-ftfy: Repair Mojibake and Unicode Text

Summary

ftfy is a Python library that repairs mojibake and other common Unicode glitches in text. It suits developers cleaning text from mixed or unreliable sources, with conservative fixes designed to avoid changing text that is already valid.

At a glance

Language
Python
License
NOASSERTION
Stars
4.1k
Forks
127
Added to OSRepos
October 21, 2025
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

ftfy repairs text that has been decoded or transformed incorrectly, especially mojibake caused by encoding mix-ups. It can also address related glitches such as HTML entities appearing in plain text and punctuation changes that interfere with decoding.

The library is for Python applications that need to clean text after it has been collected or processed. Its approach is deliberately conservative: it aims to leave sensible, correctly decoded text unchanged rather than guess at every unusual string. See the repository and documentation for details and configuration options.

Key Features

  • Repairs common UTF-8 mojibake patterns in Unicode strings.
  • Can correct multiple layers of mojibake in one pass.
  • Handles cases where curly quotes or altered spaces complicate recovery.
  • Decodes HTML entities found outside HTML, including some incorrectly capitalized forms.
  • Avoids applying a repair when text appears sensible, reducing false positives.
  • Offers text-fixing functions, configuration, and command-line usage documented by the project.

Use Cases

  • Clean text ingested from legacy databases or exports where encoding mistakes have occurred.
  • Prepare multilingual text for search, display, or analysis in a Python data pipeline.
  • Repair corrupted user-submitted or copied text before presenting it in an application.
  • Normalize scraped or aggregated content containing stray HTML entities or encoding artifacts.

Project Facts

  • Language: Python
  • License: NOASSERTION
  • Stars: 4.1k
  • Forks: 127
  • Topics: none listed
  • Archived: no

Getting Started

Install from PyPI:

pip install ftfy

See the README and documentation for usage and configuration details.

Considerations

  • The project targets Python 3.
  • Repairs rely on heuristics and are intentionally conservative. They are not a general-purpose encoding detector, so ambiguous or correctly decoded text may be left unchanged.
  • The repository metadata reports the license as NOASSERTION. Review the project’s license information before using or distributing it.

Found this useful?

Share it with someone who would like python-ftfy.

Source repository

Open the original repository on GitHub.

12 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️