tokenizers

by huggingface · Developer Tools

💥 Fast State-of-the-Art Tokenizers optimized for Research and Production

New0 ratings11,023 starsActive
Developer ToolsDeveloper ToolRustPython
Open project ↗↓ Download v0.23.2GitHub⚑ Report
tokenizers preview

About this project

Provides an implementation of today's most used tokenizers, with a focus on performance and versatility. Main features: — Train new vocabularies and tokenize, using today's most used tokenizers. — Extremely fast (both training and tokenization), thanks to the Rust implementation. Takes less than 20 seconds to tokenize a GB of text on a server's CPU. — Easy to use, but also extremely versatile. — Designed for research and production. — Normalization comes with alignments tracking. It's always possible to get the part of the original sentence that corresponds to a given token. — Does all the pre-processing: Truncate, Pad, add the special tokens your model needs. Performances Performances…

Technologies

TypeScriptPythonRustMakefileCSSgptlanguage-modelbert

Project health

Actively maintained
Last update4 days ago
Contributors143
Latest releasev0.23.2
Open issues & PRs255
LicenseApache-2.0
On GitHubsince 2019

GitHub

11,023
stars
1,189
forks
143
contributors
255
open issues & PRs
Rust
language
4 days ago
last commit
View on GitHub ↗All releases ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review tokenizers.

Built by

huggingface
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain huggingface/tokenizers? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like