Hi, I'm Liang Hsun黃亮勳
I'm an AI developer based in Taiwan. Most of my work sits at the intersection of Traditional Chinese language models and domains where being wrong is expensive — law, science, and drug discovery.
I founded Twinkle AI, an open research community started in late 2024 that builds Traditional Chinese datasets, evaluation tooling, and models with real Taiwanese cultural context — rather than treating zh-TW as a rounding error on top of Simplified Chinese.
What I work on
- Traditional Chinese LLMs. Continued pretraining and SFT on Taiwanese corpora — the Llama-3.2-Taiwan family, legal-domain variants, and a zh-TW tokenizer that compresses Traditional Chinese far better than general multilingual vocabularies.
- Data at scale. Over 160 public datasets, including
fineweb-zhtw(48M rows of filtered Traditional Chinese web text) and domain corpora for news, society, and science. - Evaluation.
tw-legal-benchmarkand Twinkle Eval — because a model that scores well on a translated English benchmark tells you almost nothing about how it handles Taiwanese statutes. - Applied ML for science. Peptide and small-molecule generation for drug discovery, including PSMA-targeted design with docking-based RL rewards.
Why this site
Papers and model cards are the polished version. This is where I put the rest — the training runs that failed, the benchmark numbers that turned out to be extraction artifacts, the ablations nobody asks for. Notes are mostly in Traditional Chinese, sometimes English.
If something here is useful, wrong, or worth arguing about, I'd like to hear it.
- Location
- Taiwan
- [email protected]
- Hugging Face
- @lianghsun
- GitHub
- @lianghsun
- Twinkle AI
- @ai-twinkle
- lianghsunhuang
- X
- @Owen_Laplace
- Kaggle
- lianghsunhuang
- Feed
- /rss.xml