Huggingface Datasets is an excellent library, but it lacks standardization, and datasets often require preprocessing work to be used interchangeably.
tasksource streamlines interchangeable datasets usage to scale evaluation or multi-task learning.
Each dataset is standardized to a MultipleChoice, Classification, or TokenClassification template with canonical fields. We focus on discriminative tasks (= with negative examples or classes) for our annotations but also provide a SequenceToSequence template. Browse the English, multilingual, and vision task catalogs for the available annotations. A preprocessing is a function that accepts a dataset and returns the standardized dataset. Preprocessing code is concise and human-readable.
pip install tasksourcefrom tasksource import list_tasks, load_task
df = list_tasks(multilingual=False) # takes some time
for id in df[df.task_type=="MultipleChoice"].id:
dataset = load_task(id) # all yielded datasets can be used interchangeablyBrowse the English, multilingual,
and vision catalogs, and feel free to request a new task.
The English catalog includes 200+ MultipleChoice tasks and 200+ Classification
tasks. Evaluation benchmarks are kept separate in
eval_only.py. Annotations excluded for duplicates,
source problems, or unsound supervision remain in
parked.py. Both record the reason for exclusion.
Use from tasksource import eval_only to access a benchmark annotation directly;
these annotations are excluded from the training catalogs.
Visual annotations are in vision_tasks.py.
Discover them with list_tasks(vision=True) and load them with
load_task(id, vision=True). Use recast="jev" or recast="instruct" to keep
images alongside the rendered decisions or prompts.
Datasets are downloaded to $HF_DATASETS_CACHE, like any Hugging Face dataset.
Ensure you have more than 100GB of space available for large multi-task runs.
Inputs are kept raw by default. When the inputs alone do not say what to predict,
an annotation carries a question ("Is this search query a well-formed question?"),
exposed as dataset.question.
load_task(id, prompted=True) appends the question to the inputs. The instruct
and typed-decision recasts use it as their instruction.
Some annotations are distributions rather than single labels: annotator votes,
rater shares, survey counts. These SoftLabeling annotations load as
probabilities with load_task(id, soft=True) (labels over options). Most also
have a hard view: their majority label on rows with clear agreement, which is
what load_task(id) and the default list_tasks() give.
list_tasks(soft=True) lists the soft views, including annotations that only
make sense as distributions, such as ProtoQA survey answers. They are listed at
the end of catalog_english.md.
Each records whether it summarizes annotator votes or mean ratings, and how many
annotators judged an item. Vote shares from three annotators are coarse
(0, 1/3, 2/3, 1); min_annotators=5 leaves them out.
list_tasks(license_use="commercial") keeps tasks whose sources allow commercial
use. The filter also accepts non-commercial, unspecified, or a list.
task_licenses() gives each task's licenses and where they come from.
Licenses are read from the Hub cards of the datasets a task loads and of their
originals, plus Data Provenance Initiative
annotations. Both are snapshotted in the package;
task_licenses(fresh=True) reads the current cards.
license_use takes the most restrictive license found. other, bare cc, and
missing licenses are unspecified. This is a best-effort filter, not legal advice.
Text encoder pretrained on tasksource reached state-of-the-art results: 🤗/deberta-v3-base-tasksource-nli
Tasksource pretraining is notably helpful for RLHF reward modeling or any kind of classification, including zero-shot. You can also find a large and a multilingual version.
The repo also contains some recasting code to convert tasksource datasets to instructions, providing one of the richest instruction-tuning datasets: 🤗/tasksource-instruct-v0
We also recast all classification tasks as natural language inference, to improve entailment-based zero-shot classification detection: 🤗/zero-shot-label-nli
Tasksource classification, multiple-choice, and vetted token tasks can be recast as
runtime-defined typed decisions (the Jev / System One request format: choice,
score, and noul questions over a state). The canonical representation keeps
the state, instructions, criteria, integer label, and textual answer separate:
🤗 tasksource/tasksource-jev-typed-decisions
The Jev build runbook covers smoke tests, resumable builds, validation, and publication.
from tasksource import load_task, render_typed_decision
dataset = load_task("glue/rte", recast="jev")
request = render_typed_decision(dataset["train"][0], model="openjev")The canonical conversion is deterministic and does not paraphrase criteria,
except that a final "all/none of the above" becomes "all/none of the other
options". Multiple-choice criteria keep every source option and are permuted per
row, seeded by task, split, and row index, so the gold slot carries no signal;
rows whose options refer to other options by position or letter keep their
order. The published 1M corpus adds explicit, deterministic, low-frequency
subrecasts for label verification (noul), criterion-order invariance, and
manually vetted instruction variation. Every row records its source, normalized
train/dev/test split, and variant. BIG-bench, MMLU, and BLiMP are excluded.
The flat rows carry group_id and question_id; render_typed_decision_group
combines related canonical decisions into one multi-question request.
Publication keeps each source-row group together under the 500k cap and uses
reviewed question and paired-field wording to reduce repeated boilerplate.
from tasksource import MultipleChoice
codah = MultipleChoice('question_propmt',choices_list='candidate_answers',
labels='correct_answer_idx',
dataset_name='codah', config_name='codah')
winogrande = MultipleChoice('sentence',['option1','option2'],'answer',
dataset_name='winogrande',config_name='winogrande_xl',
splits=['train','validation',None]) # test labels are not usable
tasks = [winogrande.load(), codah.load()]) # Aligned datasets (same columns) can be used interchangably For more details, refer to this article:
@inproceedings{sileo-2024-tasksource,
title = "tasksource: A Large Collection of {NLP} tasks with a Structured Dataset Preprocessing Framework",
author = "Sileo, Damien",
booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.lrec-main.1361",
pages = "15655--15684",
}For help integrating tasksource into your experiments, please contact damien.sileo@inria.fr.
