About

Hi, I'm Liang Hsun黃亮勳

I'm an AI developer based in Taiwan. Most of my work sits at the intersection of Traditional Chinese language models and domains where being wrong is expensive — law, science, and drug discovery.

I founded Twinkle AI, an open research community started in late 2024 that builds Traditional Chinese datasets, evaluation tooling, and models with real Taiwanese cultural context — rather than treating zh-TW as a rounding error on top of Simplified Chinese.

What I work on

  • Traditional Chinese LLMs. Continued pretraining and SFT on Taiwanese corpora — the Llama-3.2-Taiwan family, legal-domain variants, and a zh-TW tokenizer that compresses Traditional Chinese far better than general multilingual vocabularies.
  • Data at scale. Over 160 public datasets, including fineweb-zhtw (48M rows of filtered Traditional Chinese web text) and domain corpora for news, society, and science.
  • Evaluation. tw-legal-benchmark and Twinkle Eval — because a model that scores well on a translated English benchmark tells you almost nothing about how it handles Taiwanese statutes.
  • Applied ML for science. Peptide and small-molecule generation for drug discovery, including PSMA-targeted design with docking-based RL rewards.

Why this site

Papers and model cards are the polished version. This is where I put the rest — the training runs that failed, the benchmark numbers that turned out to be extraction artifacts, the ablations nobody asks for. Notes are mostly in Traditional Chinese, sometimes English.

If something here is useful, wrong, or worth arguing about, I'd like to hear it.

Location
Taiwan
Email
[email protected]
Hugging Face
@lianghsun
GitHub
@lianghsun
Twinkle AI
@ai-twinkle
LinkedIn
lianghsunhuang
X
@Owen_Laplace
Kaggle
lianghsunhuang
Feed
/rss.xml