Daniel Vila
@dvilasuero.hf.co
3.5K followers 570 following 56 posts
Everything datasets and human feedback for AI at Hugging Face. Prev: co-founder and CEO of Argilla (acquired by Hugging Face)
Posts Media Videos Starter Packs
Reposted by Daniel Vila
🚀 The open source community is unstoppable: 4M total downloads for DeepSeek models on @hf.co , with 3.2M coming from the +600 models created by the community. That's 30% more than yesterday!
Reposted by Daniel Vila
💫 Generate RAG data with the Synthetic Data Generator to improve your RAG system!

1️⃣ Generate from your documents, dataset, or dataset description.
2️⃣ Configure it.
3️⃣ Generate the synthetic dataset.
4️⃣ Fine-tune the retrieval and reranking models.
5️⃣ Build a RAG pipeline.
Reposted by Daniel Vila
New chapter in the Hugging Face NLP course! 🤗 🚀

We've added a new chapter about the very basics of Argilla to the Hugging Face NLP course. Learn how to set up an Argilla instance, load & annotate datasets, and export them to the Hub. 

Any feedback for improvements welcome!
Reposted by Daniel Vila
🎉 50,000+ annotations reached! The FineWeb2-C community is helping build better language models on annotation at a time.

📊 Current stats:
- 115 languages represented
- 419 amazing contributors
- 24 languages with complete datasets

But we're not done yet! 🧵
Reposted by Daniel Vila
High-quality data for fine-tuning language models for free and at the click of a button!

Prompt and wait for your dataset to push to Argilla or the Hub
Evaluate, review and fine-tune a model.

Blog:
Fine-tune a SmolLM on domain-specific synthetic data from a LLM
A Blog post by David Berenstein on Hugging Face
buff.ly
Reposted by Daniel Vila
Was 2024 the year of datasets? Is 2025 the year for community-built datasets?

It's exciting to see the progress of many languages in FineWeb-C:
- Total annotations submitted: 41,577
- Languages with annotations: 106
- Total contributors: 363
Reposted by Daniel Vila
The finish line is near! We're building FineWeb-Edu for many languages and need your help 🤗

Many FineWeb-C languages are close to 1,000 annotations!

Assamese is 99.4% done, French needs 64 more annotations, Tamil: 216.

Please help us reach the goal: huggingface.co/spaces/data-...
💥 Ending 2024: A full data annotation journey on the Hugging Face Hub—from raw data to training-ready datasets!

With Argilla 2.6.0, push your data to the Hub from the UI

Let’s make 2025 the year anyone can build more transparent and accountable AI—no coding or model skills needed.
Reposted by Daniel Vila
🚀 Argilla v2.6.0 is here! 🎉

Let me show you how EASY it is to export your annotated datasets from Argilla to the Hugging Face Hub. 🤩

Take a look to this quick demo 👇

💁‍♂️ More info about the release at github.com/argilla-io/a...

#AI #MachineLearning #OpenSource #DataScience #HuggingFace #Argilla
Reposted by Daniel Vila
🔥 We got great feedback on this: "Synthetic Data Generator"

A no-code tool to create datasets with LLMs, making it a breeze, allowing ANYONE to create datasets and models in minutes and without any code.

Blog: https://buff.ly/4gybyoT
GitHub: https://buff.ly/49IDSmd
Space: https://buff.ly/3Y1S99z
Introducing the Synthetic Data Generator - Build Datasets with Natural Language
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
buff.ly
Reposted by Daniel Vila
Well, around 10 percent of the initial goal is complete, and so far, it's been quite a one-man army effort. We're still in the hunt for more people to join and contribute to this open-source initiative.

@hf.co

data-is-better-together-fineweb-c.hf.space/share-your-p...
tam - தமிழ் - Tamil
Join and contribute to the dataset tam - தமிழ் - Tamil
data-is-better-together-fineweb-c.hf.space
Reposted by Daniel Vila
I've been building a small library for working with prompt templates on the @huggingface.bsky.social Hub: `pip install prompt-templates`. Motivation:

The community currently shares prompt templates in a wide variety of formats: in datasets, in model cards, as strings in .py files, as .txt/... 🧵
Help shape the future of multilingual Open Source AI!

Join the FineWeb 2 Community Annotation Sprint to create an open training dataset with full transparency and human validation in many languages.

Review datasets in your language and help identify the best sources for training.
Reposted by Daniel Vila
✨ Argilla 2.5.0 is live and it comes with webhook listener support to supercharge your workflows! 🚀

#AI #MachineLearning #Webhooks #TechUpdate
Reposted by Daniel Vila
👐 Open Image Preferences is an Apache 2.0 licensed dataset for text-to-image generation by the @hf.co community. This dataset contains 10K text-to-image preference pairs across image generation categories, using different model families and prompt complexities.

Blog: huggingface.co/blog/image-p...
Open Preference Dataset for Text-to-Image Generation by the 🤗 Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
Reposted by Daniel Vila
Open Image Preferences released! 🚀

- Open-source dataset for text2image
- 10K samples manually evaluated by the HF community.
- Binarized format for SFT, DPO, or ORPO.

It comes with a nice blog post explaining the steps to pre-process and generate the data, along with the results.
I'd love to yes!!