From 169114176573bf8c4fff62a5ed6e20b5b69d6096 Mon Sep 17 00:00:00 2001 From: PENG Bo <33809201+BlinkDL@users.noreply.github.com> Date: Mon, 16 May 2022 22:59:46 +0800 Subject: [PATCH] Update README.md --- README.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/README.md b/README.md index 6352bea..3acbd24 100644 --- a/README.md +++ b/README.md @@ -27,6 +27,8 @@ https://github.com/BlinkDL/RWKV-LM/tree/main/RWKV-v2-RNN Note: For fine-tuning the Pile model, change 1e-15 to 1e-9 in https://github.com/BlinkDL/RWKV-LM/blob/main/RWKV-v2-RNN/src/model.py and https://github.com/BlinkDL/RWKV-LM/blob/main/RWKV-v2-RNN/src/model_run.py and probably you need other changes as well. You can compare the output with the latest code ( https://github.com/BlinkDL/RWKV-v2-RNN-Pile ) to verify it. +I usually fine-tune with 4e-5 lr, and decay to 1e-5 when it plateaus. + ## How it works RWKV is inspired by Apple's AFT (https://arxiv.org/abs/2105.14103).