Pre-training a Bengali BERT From Scratch on a 16GB M1 Mac — and Onto an IndicGLUE Leaderboard
A 7.5M-parameter BERT pre-trained from random weights on Bengali Wikipedia in 28 laptop-hours, then fine-tuned to 86.5% on IndicGLUE sna.bn — above published mBERT and IndicBERT.
![86.5% on IndicGLUE sna.bn, above mBERT and IndicBERT — from 7.5M parameters trained 28 hours on one 16GB M1. On the right, the six-layer encoder stack: a [CLS] A [SEP] B [SEP] pair goes in at the top, a pooled [CLS] classifier reads it out at the bottom.](/_astro/bert_cover.ChghBB-Z_Z16YI7L.webp)
