pytorch autograd 自动微分与梯度更新

yichudu

已于 2023-01-10 11:52:40 修改

阅读量1.3k

点赞数 1

CC 4.0 BY-SA版权

分类专栏： torch 数学概率统计最优化文章标签： pytorch 深度学习 python

于 2022-09-20 17:55:22 首次发布

天天开心

本文链接：https://blog.csdn.net/chuchus/article/details/126958252

数学概率统计最优化同时被 2 个专栏收录

44 篇文章

订阅专栏

torch

7 篇文章

订阅专栏

本文详细介绍了PyTorch中自动微分的工作原理及应用。通过实例演示了如何利用链式法则进行复合函数求导，并展示了PyTorch中张量的梯度计算流程。此外，还介绍了梯度更新的方法及其相关API。

摘要生成于 C知道，由 DeepSeek-R1 满血版支持，前往体验 >

自动微分

原理

pytorch 内置了常见 tensor 操作的求导解析解. 从 loss 到 parameter 是若干个 op 叠加起来的复合函数, 所以用链式法则逐个计算.
tensor.grad_fn 记录了一个 tensor 是由何种运算产出的, 以及相应的求导解析解. 注意并不是根据 $yx′=ΔyΔxy'_x=\frac {\Delta y} {\Delta x}$ 的定义去计算数值解.

例子

令 $x_1=2,x_2=2$
$u=(u_1,u_2)=f_1(x)=(4x_1,4x_2)$
$y=f2(u)=(u12+u22)12y=f_2(u)=(u_1^2+u_2^2)^{\frac12}$ ,
求 y 对 x1 的偏导数.

手算偏导

使用链式法则作复合函数的求导.
$∂y∂x1=∂y∂u1⋅∂u1∂x1(1)\frac{\partial y}{\partial x_1} = \frac{\partial y}{\partial u_1} \cdot \frac{\partial u_1}{\partial x_1} \tag1$

分别计算两项各自的导数:
$∂y∂u1=∂(u12+u22)12∂u1=12×(u12+u22)−12×2u1=12×164+64×2×8=0.7071(2)\frac{\partial y}{\partial u_1}=\frac{{\partial}(u_1^2+u_2^2)^{\frac12}}{{\partial u_1}}\\ =\frac12 \times (u_1^2+u_2^2)^{-\frac12}\times 2u_1\\ =\frac12\times\frac1{\sqrt{64+64}}\times 2\times8\\ =0.7071 \\ \tag2$
注意因为有 $u_1^2$ 的存在, 这里其实也是一个复合函数, 都用到了 $x^a)'=ax^{a-1}$ 的求导公式.

$∂u1∂x1=4(3)\frac{\partial u_1}{\partial x_1} =4 \tag3$

将 (2)(3)的结果代入式(1), 有
$\frac{\partial y}{\partial u_1} \cdot \frac{\partial u_1}{\partial x_1}\\ =0.7071\times 4\\ =2.8284 \tag4$

代码比对

import torch

x = torch.tensor([2, 2], dtype=torch.float)  # input tensor
x.requires_grad = True
u = 4 * x
y: torch.Tensor = u.norm()
print('x.grad_fn', x.grad_fn)
print('y.grad_fn', y.grad_fn)
print('u.grad_fn', u.grad_fn)
print(f'before y.backward(), x.grad = {x.grad}, u.grad = {u.grad}')
y.backward()
print(f'after y.backward(), x.grad = {x.grad}, u.grad = {u.grad}')
"""
x.grad_fn None
y.grad_fn <NormBackward1 object at 0x0000027620F89A90>
u.grad_fn <MulBackward0 object at 0x0000027620F89A90>
before y.backward(), x.grad = None, u.grad = None
after y.backward(), x.grad = tensor([2.8284, 2.8284]), u.grad = None
D:\code_study\torch_study\test\auto_grad_test.py:10: UserWarning: The .grad attribute of a Tensor that is not a leaf Tensor is being accessed. Its .grad attribute won't be populated during autograd.backward(). If you indeed want the .grad field to be populated for a non-leaf Tensor, use .retain_grad() on the non-leaf Tensor. If you access the non-leaf Tensor by mistake, make sure you access the leaf Tensor instead. See github.com/pytorch/pytorch/pull/30531 for more informations. (Triggered internally at C:\actions-runner\_work\pytorch\pytorch\builder\windows\pytorch\build\aten\src\ATen/core/TensorBody.h:485.)
  print(f'before y.backward(), x.grad = {x.grad}, u.grad = {u.grad}')
"""

可以清晰看到 y对x1的偏导为 2.8284, 与手算结果一致.
有个 warning 信息, 是说 tensor u 不是叶子结点, 所以.grad attribute 不会被自动计算. 如果硬要算也可以, 调用 u.retain_grad() 即可.

不可微的 op 怎么搞?

一些分段函数带来的第一类间断点等.
todo

梯度更新

手算tensor迭代

使用最简单的 SGD, 步长为 0.1, 那么一个 step 之后, x 新的值为
x=x+(-1)*gradient*learning_rate, 代入得 x=2-0.1*2.8284=1.7172.

代码比对

import torch

x = torch.tensor([2, 2], dtype=torch.float)  # input tensor
x.requires_grad = True

u: torch.Tensor = 4 * x
y: torch.Tensor = u.norm()
print('y.grad_fn = ', y.grad_fn)
print('u.grad_fn = ', u.grad_fn)
print('x.grad_fn = ', x.grad_fn)

loss = y
optimizer = torch.optim.SGD(params=[x], lr=0.1)

u.retain_grad()
optimizer.zero_grad()

print(f'before loss.backward(), x.grad = {x.grad}, u.grad = {u.grad}')
loss.backward()
print(f'after loss.backward(), x.grad = {x.grad}, u.grad = {u.grad}')

print(f'before optimizer.step(), x = {x}')
optimizer.step()
print(f'after optimizer.step(), x = {x}')

"""
y.grad_fn <CopyBackwards object at 0x0000022F2CDA1BE0>
u.grad_fn <MulBackward0 object at 0x0000022F2CDA1BE0>
x.grad_fn None

before loss.backward(), x.grad = None, u.grad = None
after loss.backward(), x.grad = tensor([2.8284, 2.8284]), u.grad = tensor([0.7071, 0.7071])

before optimizer.step(), x = tensor([2., 2.], requires_grad=True)
after optimizer.step(), x = tensor([1.7172, 1.7172], requires_grad=True)
"""