Hugging Face
Models
Datasets
Spaces
Posts
Docs
Enterprise
Pricing
Log In
Sign Up
91.9
TFLOPS
3
4
Kishore Kashyap
kishorekashyap
Follow
oliviermills's profile picture
davanstrien's profile picture
2 followers
·
2 following
kishorekashyap
AI & ML interests
NLP, ML, AI
Recent Activity
new
activity
about 3 hours ago
HuggingFaceFW/fineweb-2:
Synthetic Data Generator
reacted
to
davanstrien
's
post
with ❤️
about 3 hours ago
The https://huggingface.co./datasets/data-is-better-together/fineweb-c dataset is growing! This week a few more languages have got 1,000 annotations for the educational quality of data from https://huggingface.co./datasets/HuggingFaceFW/fineweb-2. Why should you care? The quality of pre-training data can have a big impact on the performance of downstream language models trained on that data (https://huggingface.co./spaces/HuggingFaceFW/blogpost-fineweb-v1). Being able to filter by educational quality is on way of improving the quality of the data you use for training an LLM. Very importantly this approach can also reduce the amount of data needed for pertaining. Why not use an LLM? LLMs can be used to annotate educational quality for a subset of data. This data can then be used to train a smaller encoder only model to label the full dataset. However, this may not work well for languages outside of english. This is where fineweb-c (community) comes in. The community is annotating the educational quality of fineweb2 data. Currently 114 languages have some annotations. These annotations will enable a number of things: - Evaluate whether an LLM can label the educational quality for texts in that language well - Directly be used for training quality classifiers - Help discover other rules and huerisitcs for refining fineweb2 further for different languages. This week the following languages where done: Swedish thanks to: @Lauler @AntonVic @ohallstrom @bjarlestam @menbom @Ekgren @apsod Ukrainian thanks to: @hannayukhymenko @robinhad @realPivo @RabotiahovDmytro @reciprocate Assamese thanks to: @moyoor97 @Arpanjyoti @nawaf-helmi123 @pahigogoi1 @aelhence @kishorekashyap Want to learn more: https://huggingface.co./blog/davanstrien/fineweb2-community Contribute yourself here: https://huggingface.co./spaces/data-is-better-together/fineweb-c
new
activity
about 1 month ago
TWO/sutra-mlt256-v2:
TWO/sutra-mlt256-v2 does not appear to have a file named config.json
View all activity
Organizations
kishorekashyap
's activity
All
Models
Datasets
Spaces
Papers
Collections
Community
Posts
Upvotes
Likes
liked
a dataset
10 months ago
google/fleurs
Updated
Aug 25, 2024
•
17.1k
•
262
liked
a Space
10 months ago
Running
58
⚔
Tokenizer Arena
Compare different tokenizers in char-level and byte-level.
liked
2 models
10 months ago
google/gemma-7b
Text Generation
•
Updated
Jun 27, 2024
•
37.3k
•
•
3.09k
MaLA-LM/mala-500-10b-v1
Text Generation
•
Updated
Apr 3, 2024
•
57