Sapin-sapin AI — Sariling Ai PINas 🇵🇭
The Philippines speaks more than 180 languages, the support for the
national and regional lingua franca is lackluster and almost none of
the others are meaningfully represented in today's AI models.
Sapin-sapin (Sariling Ai PINas) exists to change that: we build the open
foundations — data, benchmarks, and models — for Philippine-language AI.
Like the layered kakanin we're named after, the work stacks: raw archives and
recordings become cleaned corpora, corpora become training sets, training sets
become models anyone can build on. The initiative grew out of the
Philippine AI Report and is featured in
Featherless.ai's
No Language Forgotten
case study.
🎙️ Try the models
halohalo-dashboard —
transcribe, synthesize, or convert a voice in ten Philippine languages,
straight from your microphone or the preloaded clips. The same Space carries a
live view of everything in this org, queried from the Hub API.
🗂️ Datasets
Speech
- filipinospeechcorpus — the UP-DSP Filipino Speech Corpus as sentence-level 16 kHz segments, Common-Voice-schema, ASR/TTS-ready
- pld — the UP-DSP Philippine Language Dataset: 334k utterances, 448 hours, 980 speakers across ten languages
- halo-livestream — diarized livestream speech with forced alignment and QC gating, in ASR (16 kHz) and TTS (24 kHz) variants
Text
- BantayWika — Philippine literary and reference corpora (Filipiniana, Gutenberg, newspapers, Palito, FilNet, ISIP) in FineWeb-compatible JSONL
- halohalo — combined web corpus from CommonCrawl, cleaned and deduplicated, with per-language layers: halo-tgl (Tagalog), halo-hil (Hiligaynon), halo-bcl (Bikol)
🤖 Models
- LLMs adapted to Philippine news and languages: llama31-8b-balitanlp-cpt, llama31-8b-balitanlp-IT, gpt-oss-20b-balitanlp-cpt, qwen3vl-balitanlp-news-writer, bikoLLM
- Speech — 23 models across ten Philippine languages, finetuned on the
Philippine Language Dataset
(Bikol, Cebuano, Filipino, Hiligaynon, Ilocano, Kapampangan, Pangasinan,
Tausug, Waray, Philippine English). For most of these languages they are the
first public models of their kind:
- Speech recognition —
whisper-small-pld-<lang>, one per language
- Speech synthesis —
speecht5_tts-pld-<lang>, one per language
- Voice conversion — speecht5_vc-pld, any-to-any across all ten
- On the Filipino Speech Corpus: speecht5_tts-fsc, whisper-small-fsc
🤝 It takes a village
"It is not easy — it takes a village for us to succeed. That is precisely why
I have made community collaboration the foundation of every single
initiative."
— Tim Santos, founder
If you work on Philippine languages — as a researcher, annotator, native
speaker, or engineer — we want to hear from you. Open an issue or discussion on
any repo here.