jieba: Segment Chinese Text in Python

jieba: Segment Chinese Text in Python

Summary

jieba is a Python library for segmenting Chinese text into words, including text that does not use spaces as word boundaries. It offers configurable segmentation modes, custom dictionaries, part-of-speech tagging, and keyword extraction.

At a glance

Language
Python
License
MIT
Stars
35.2k
Forks
6.7k
Added to OSRepos
March 31, 2026
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

jieba is a Python library for Chinese word segmentation. It addresses the challenge of finding word boundaries in Chinese text, supporting workflows such as text analysis and search indexing.

Its dictionary-and-probability-based approach offers several segmentation modes and lets developers adapt vocabulary and tokenization behavior. The repository also includes tools for part-of-speech tagging, keyword extraction, and command-line use.

Key Features

  • Accurate mode aims for a single, precise segmentation; full mode returns possible dictionary matches; search mode further splits longer words for retrieval.
  • Uses a word-frequency dictionary, dynamic programming, and an HMM with Viterbi decoding for unknown words.
  • Supports custom dictionaries and runtime vocabulary adjustments.
  • Provides part-of-speech tagging and keyword extraction with TF-IDF or TextRank.
  • Can return token offsets and offers a command-line interface.
  • Includes optional PaddlePaddle-based segmentation and tagging, which requires installing paddlepaddle-tiny.

Use Cases

  • NLP practitioners can prepare Chinese text for downstream analysis when whitespace does not mark word boundaries.
  • Search developers can use fine-grained search-mode tokens to build or improve text indexing.
  • Teams working with domain-specific terms can load a custom dictionary to guide segmentation.
  • Analysts can extract keywords or attach part-of-speech labels as part of text-processing workflows.

Project Facts

  • Language: Python
  • License: MIT
  • Stars: 35.2k
  • Forks: 6.7k
  • Topics: none listed
  • Archived: false

Getting Started

Install from PyPI and import the package:

pip install jieba
import jieba
words = jieba.cut("?????????")

See the README for usage details, configuration, and optional dependencies.

Considerations

  • Segmentation quality depends on the dictionary and word frequencies. Domain terms may need a custom dictionary or frequency adjustments.
  • HMM-based unknown-word discovery can affect results. The README documents how to disable it when needed.
  • Paddle mode requires the separate paddlepaddle-tiny dependency.
  • The repository reports 700 open issues, so review current project activity and open issues when evaluating it for a new dependency.

Source repository

Open the original repository on GitHub.

17 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️