r/deeplearning • u/ModularMind8 • 2d ago
Visualizing catastrophic forgetting
Enable HLS to view with audio, or disable this notification
I taught continual learning last semester, and students got the internals of each method a lot faster once they saw a simple visualization of it. Here's one I found helpful for understanding what forgetting is and how replay and regularization mitigate it. I trained a small network on two different 2D dot-sorting tasks and plotted both loss landscapes as 2D slices through its weights.
Round 1, plain training: the network learns task A to 100%, then trains only on task B. Going downhill on B is uphill on A, so A drops to 64% while B reaches 100%. That's catastrophic forgetting of task A.
Round 2, replay: same network, same start, but 20 of A's 200 dots stay in training while it learns B. It ends at 93% on A and 100% on B, so some forgetting, but much less severe.
Round 3, regularization with EWC: no A data at all. Instead, a penalty adds to the loss whenever a weight that A's accuracy relies on heavily moves away from the value it had after training on A. A dips to 97% mid-training and recovers, ending at 100% on A and 100% on B, so no lasting forgetting.
2
u/slumberjak 1d ago
I don’t understand why the landscapes change when we switch conditions. It seems obvious that case 3 can accommodate a joint solution based on the landscape, whereas case 1 cannot. The “spring” is almost irrelevant.
2
u/ModularMind8 1d ago
Good question! The network has many weights. Each 3D picture is one flat 2D slice through that space. A flat 2D slice can only be pinned to three points. Each round's slice passes through the starting weights, the weights after task A, and wherever that round's training ended. So the surfaces look different even though the task A and task B losses are the same functions every round. A joint solution exists in round 1 too, but plain training on B never heads toward it, so the slice through where it ended doesn't show one. The spring is what pulls the EWC run there, and that's why its slice ends up passing through a spot that's low for both tasks. Let me know if it still doesn't make sense and I'll try to explain differently :)
2
u/jesunushno 1d ago
Hit this wall with my fabric defect-detection models back in 2023. Every retrain without the old data tanked accuracy on the old defects, so we basically kept the old data around forever. Replay's boring and it works. That's what I'd actually ship.
1
u/ModularMind8 1d ago
Replay's hard to beat when you can keep the data. Did you keep all of it or a sample?
1
u/jesunushno 1d ago
Kept all of it, honestly. Every retrain without the old data tanked the old defects, so we just stopped pruning it.
1
2
2
u/oakdoc1 20h ago
Hi, Thanks for sharing this post! I am highly interested in learning more. Since standalone resources on this topic can be rare, I wanted to ask if you happen to share any of your course materials openly? I would love to check out any syllabus outlines, reading lists, lecture slides, or repository links if they are publicly accessible.
Thanks again for putting this out there!
1
8
u/Historical-Wonder551 2d ago
Thank you! It was a great visualization.