From 8d63423a56e43b7fa75163daf544c244911778fe Mon Sep 17 00:00:00 2001 From: PENG Bo <33809201+BlinkDL@users.noreply.github.com> Date: Mon, 20 Jun 2022 18:04:40 +0800 Subject: [PATCH] Update README.md --- README.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/README.md b/README.md index 088e06c..d3ec092 100644 --- a/README.md +++ b/README.md @@ -49,6 +49,8 @@ My LR schedule for the L24-D1024 RWKV-2: Fixing NaN or loss spikes: load a previous checkpoint, decrease LR a bit. I find you can decrease the LR faster than GPT, and eventually to 1/50 of LR_max. +**UPDATE: Search for "RWKV v2+" here and change RWKV-2 to PreLN to make it more stable.** + Fine-tuning: see https://github.com/BlinkDL/RWKV-v2-RNN-Pile. ## How it works