Resource of free step by step video how to guides to get you started with machine learning.
Thursday, April 23, 2020
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
This video explores the T5 large-scale study on Transfer Learning. This paper takes apart many different factors of the Pre-Training then Fine-Tuning pipeline for NLP. This involves Auto-Regressive Language Modeling vs. BERT-Style Masked Language Modeling and XLNet-style shuffling, as well as the impact of dataset composition, size, and how to best use more computation. Thanks for watching and please check out Machine Learning Street Talk where Tim Scarfe, Yannic Kilcher and I discuss this paper! Machine Learning Street Talk: https://www.youtube.com/channel/UCMLtBahI5DMrt0NPvDSoIRQ Paper Links: T5: https://ift.tt/2pcuaXx Google AI Blog Post on T5: https://ift.tt/2SV4VF9 Train Large, Then Compress: https://ift.tt/3awfC74 Scaling Laws for Neural Language Models: https://ift.tt/2yzOOVY The Illustrated Transformer: https://ift.tt/2NLJXmf ELECTRA: https://ift.tt/2RZsM5S Transformer-XL: https://ift.tt/2LIaXXb Reformer: The Efficient Transformer: https://ift.tt/378kuhh The Evolved Transformer: https://ift.tt/2IAdYFw DistilBERT: https://ift.tt/2Y2cZa2 How to generate text (HIGHLY RECOMMEND): https://ift.tt/3d9QC7P Tokenizers: https://ift.tt/2vpu7Kx Thanks for watching! Please Subscribe!
Subscribe to:
Post Comments (Atom)
-
Using GPUs in TensorFlow, TensorBoard in notebooks, finding new datasets, & more! (#AskTensorFlow) [Collection] In a special live ep...
-
Quick answer: Qwen 3 32B needs about 22.7 GB of VRAM at Q4_K_M with a 4k-token context — a used 24 GB card like the RTX 3090 is the minimu...
-
Inside TensorFlow: Summaries and TensorBoard [Collection] Take an inside look into the TensorFlow team’s own internal training sessions-...
-
Hey everyone This is Ujjwal kapil B.tech 2nd year Ai and Ml In this video we covered one of the most important question "COLLEGE KAB O...
-
Quick answer: DeepSeek V4 Flash is a Mixture-of-Experts model: all 284 billion parameters must be stored, so combined RAM plus VRAM — not ...
-
This video is a crash course on understanding how finetuning on LLM models can be performed uing QLORA,LORA, Quantization using LLama2, Grad...
-
In this video, you'll learn how to use machine learning, computer vision and deep learning to create a football analysis system. This pr...
-
Quick answer: For most learners, the best single machine learning book is Hands-On Machine Learning with Scikit-Learn, Keras & Tensor...
-
#minecraft #neuralnetwork #backpropagation I built an analog neural network in vanilla Minecraft without any mods or command blocks. The n...
-
Quick answer: For an 8B model at Q4, local electricity costs roughly $0.05 per million output tokens — against about $0.53 per million ble...
No comments:
Post a Comment