残差网络（Residual Networks, ResNets）

1. 什么是残差

　　“残差在数理统计中是指实际观察值与估计值（拟合值）之间的差。”“如果回归模型正确的话，我们可以将残差看作误差的观测值。”

　　更准确地，假设我们想要找一个

　　即使

2. 什么是残差网络（Residual Networks，ResNets）？

　　在了解残差网络之前，先了解下面这个问题。

　　Q1：神经网络越深越好吗？（Deeper is better？）

　　A1：如图 1 所示，在训练集上，传统神经网络越深效果不一定越好。而 Deep Residual Learning for Image Recognition 这篇论文认为，理论上，可以训练一个 shallower 网络，然后在这个训练好的 shallower 网络上堆几层 identity mapping（恒等映射）的层，即输出等于输入的层，构建出一个 deeper 网络。这两个网络（shallower 和 deeper）得到的结果应该是一模一样的，因为堆上去的层都是 identity mapping。这样可以得出一个结论：理论上，在训练集上，Deeper 不应该比 shallower 差，即越深的网络不会比浅层的网络效果差。但为什么会出现图 1 这样的情况呢，随着层数的增多，训练集上的效果变差？这被称为退化问题（degradation problem），原因是随着网络越来越深，训练变得原来越难，网络的优化变得越来越难。理论上，越深的网络，效果应该更好；但实际上，由于训练难度，过深的网络会产生退化问题，效果反而不如相对较浅的网络。而残差网络就可以解决这个问题的，残差网络越深，训练集上的效果会越好。（测试集上的效果可能涉及过拟合问题。过拟合问题指的是测试集上的效果和训练集上的效果之间有差距。）

图 1 不同深度的传统神经网络效果对比图

（“plain” network指的是没有使用 shortcut connection 的网络）

　　残差网络通过加入 shortcut connections，变得更加容易被优化。包含一个 shortcut connection 的几层网络被称为一个残差块（residual block），如图 2 所示。（shortcut connection，即图 2 右侧从

图 2 残差块

　　2.1 残差块（residual block）

　　如图 2 所示，

　　当没有 shortcut connection（即图 2 右侧从

　　2.2 残差网络举例

　　图 3 最右侧就是就是一个残差网络。34-layer 表示含可训练参数的层数为34层，池化层不含可训练参数。图 3 右侧所示的残差网络和中间部分的 plain network 唯一的区别就是 shortcut connections。这两个网络都是当 feature map 减半时，filter 的个数翻倍，这样保证了每一层的计算复杂度一致。

　　ResNet 因为使用 identity mapping，在 shortcut connections 上没有参数，所以图 3 中 plain network 和 residual network 的计算复杂度都是一样的，都是 3.6 billion FLOPs.

图 3 VGG-19、plain network、ResNet

　　残差网络可以不是卷积神经网络，用全连接层也可以。当然，残差网络在被提出的论文中是用来处理图像识别问题。

　　2.3 为什么残差网络会work？

　　我们给一个网络不论在中间还是末尾加上一个残差块，并给残差块中的 weights 加上 L2 regularization（weight decay），这样图 1 中

　　"The main reason the residual network works is that it‘s so easy for these extra layers to learn the identity function that you‘re kind of guaranteed that it doesn‘t hurt performance. And then lot of time you maybe get lucky and even helps performance, or at least is easier to go from a decent baseline of not hurting performance, and then creating the same can only improve the solution from there."

References

残差百度百科

Residual (numerical analysis) - Wikipedia

Course 4 Convolutional Neural Networks by Andrew Ng

He, K., Zhang, X., Ren, S., & Sun, J. (2015). Deep Residual Learning for Image Recognition, 1–9. Retrieved from http://arxiv.org/abs/1512.03385

残差网络（Residual Networks, ResNets）

标签：举例 tin sid entity detail 回归 tee 过拟合 bsp

原文地址：https://www.cnblogs.com/hdhdfgdsfg/p/12078286.html

踩

(0)

评论一句话评论（0）