datasets
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
About this project
🤗 Datasets is a lightweight library providing two main features: — one-line dataloaders for many public datasets: one-liners to download and pre-process any of the major public datasets (image datasets, audio datasets, text datasets in 467 languages and dialects, 3D medical images, video datasets, agent traces, etc.) provided on the HuggingFace Datasets Hub. With a simple command like squaddataset = loaddataset("rajpurkar/squad"), get any of these datasets ready to use in a dataloader for training/evaluating a ML model (Numpy/Pandas/PyTorch/TensorFlow/JAX/Polars), — efficient data pre-processing: simple, fast and reproducible data pre-processing for the public datasets as well as your own…
Technologies
Project health
GitHub
Reviews
Built by
Maintain huggingface/datasets? Claiming verifies admin access through your GitHub account and gives you control of this listing.