Home/The papers that got us here/BERT: Pre-training of Deep Bidirectional TransformersItemBERT: Pre-training of Deep Bidirectional Transformersin The papers that got us here by TheLysts PlatformLike this itemFollow TheLysts PlatformOpen in appBERT: Pre-training of Deep Bidirectional Transformers on “The papers that got us here”, a list by TheLysts Platform on TheLysts.DetailsCommentMade pretrain-then-finetune the standard recipe for NLP and showed transfer learning works for language.PreviousAttention Is All You NeedNextLanguage Models are Few-Shot Learners (GPT-3)Related itemsComputing Machinery and IntelligenceLearning Representations by Back-propagating ErrorsLong Short-Term MemoryImageNet Classification with Deep Convolutional Neural Networks (AlexNet)Report