Huang Liang Hsun黃亮勳
AI developer working on Traditional Chinese language models, evaluation, and applied ML for law and science.
打造繁體中文語言模型、評測基準,以及把機器學習用在法律與科學領域。
Writing
All posts →Selected work
Everything →TwinkleTokenizer
A 201K byte-level BPE vocabulary trained from scratch on Traditional Chinese — 38% lower tokens/char than Qwen3.8-27B on tw-tokenizer-bench, with a smaller vocabulary.
201,069 vocab · NFC · Apache-2.0 ModelLlama-3.2-Taiwan-3B-Instruct
Instruction-tuned Llama 3.2 adapted to Taiwanese Traditional Chinese and local context.
71 likes on the Hub Pretraining corpusfineweb-zhtw
A large-scale Traditional Chinese web corpus, filtered and deduplicated for LLM pretraining.
48.1M rows Benchmarktw-tokenizer-bench
Traditional Chinese tokenizer benchmark across five formal-register domains, with fixed character windows and training-contamination shards excluded.
24,294 samples · 18.4M chars PlatformHub
Community hub for Traditional Chinese AI datasets and tooling — hub.twinkleai.tw.
186 stars Translationrlhf-book-zh-tw
Traditional Chinese translation of the RLHF Book, with runnable labs.
177 stars