Resource of free step by step video how to guides to get you started with machine learning.
Thursday, April 23, 2020
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
This video explores the T5 large-scale study on Transfer Learning. This paper takes apart many different factors of the Pre-Training then Fine-Tuning pipeline for NLP. This involves Auto-Regressive Language Modeling vs. BERT-Style Masked Language Modeling and XLNet-style shuffling, as well as the impact of dataset composition, size, and how to best use more computation. Thanks for watching and please check out Machine Learning Street Talk where Tim Scarfe, Yannic Kilcher and I discuss this paper! Machine Learning Street Talk: https://www.youtube.com/channel/UCMLtBahI5DMrt0NPvDSoIRQ Paper Links: T5: https://ift.tt/2pcuaXx Google AI Blog Post on T5: https://ift.tt/2SV4VF9 Train Large, Then Compress: https://ift.tt/3awfC74 Scaling Laws for Neural Language Models: https://ift.tt/2yzOOVY The Illustrated Transformer: https://ift.tt/2NLJXmf ELECTRA: https://ift.tt/2RZsM5S Transformer-XL: https://ift.tt/2LIaXXb Reformer: The Efficient Transformer: https://ift.tt/378kuhh The Evolved Transformer: https://ift.tt/2IAdYFw DistilBERT: https://ift.tt/2Y2cZa2 How to generate text (HIGHLY RECOMMEND): https://ift.tt/3d9QC7P Tokenizers: https://ift.tt/2vpu7Kx Thanks for watching! Please Subscribe!
Labels:
Video
Published by Free artificial intelligence and machine learning video tutorial resource
Related video tutorials
Subscribe to:
Post Comments (Atom)
Most watched
-
Quick answer: Qwen 3 32B needs about 22.7 GB of VRAM at Q4_K_M with a 4k-token context — a used 24 GB card like the RTX 3090 is the minimu...
-
Using GPUs in TensorFlow, TensorBoard in notebooks, finding new datasets, & more! (#AskTensorFlow) [Collection] In a special live ep...
-
Inside TensorFlow: Summaries and TensorBoard [Collection] Take an inside look into the TensorFlow team’s own internal training sessions-...
-
Quick answer: DeepSeek V4 Flash is a Mixture-of-Experts model: all 284 billion parameters must be stored, so combined RAM plus VRAM — not ...
-
Quick answer: For an 8B model at Q4, local electricity costs roughly $0.05 per million output tokens — against about $0.53 per million ble...
-
Hey everyone This is Ujjwal kapil B.tech 2nd year Ai and Ml In this video we covered one of the most important question "COLLEGE KAB O...
-
Quick answer: For machine-learning work, buy an external SSD over a spinning hard drive, pick USB 3.2 Gen 2 (10 Gbps) or faster, and size f...
-
Quick answer: Yes, for most budget builders: a used RTX 3090 is still the cheapest route to 24 GB of VRAM with full CUDA support. Capacity...
-
This video is a crash course on understanding how finetuning on LLM models can be performed uing QLORA,LORA, Quantization using LLama2, Grad...
-
In this video, you'll learn how to use machine learning, computer vision and deep learning to create a football analysis system. This pr...
No comments:
Post a Comment