实现 Multi-Head Attention 算法。
给定输入序列 X(L×d)、头数 h,以及投影矩阵 Wq,Wk,Wv,Wo(各 d×d),按以下步骤计算输出:
第一行包含三个整数 L,d,h(1≤L≤8,2≤d≤8,1≤h≤d,dmodh=0),分别表示序列长度、模型维度、头数。
接下来 L 行,每行 d 个浮点数,表示输入序列 X。
接下来 d 行,每行 d 个浮点数,表示 Wq。
接下来 d 行,每行 d 个浮点数,表示 Wk。
接下来 d 行,每行 d 个浮点数,表示 Wv。
接下来 d 行,每行 d 个浮点数,表示 Wo。
输出 L 行,每行 d 个浮点数,表示 Multi-Head Attention 的输出矩阵,保留 4 位小数。
输入
2 4 2
1.0 0.0 1.0 0.0
0.0 1.0 0.0 1.0
1.0 0.0 0.0 0.0
0.0 1.0 0.0 0.0
0.0 0.0 1.0 0.0
0.0 0.0 0.0 1.0
1.0 0.0 0.0 0.0
0.0 1.0 0.0 0.0
0.0 0.0 1.0 0.0
0.0 0.0 0.0 1.0
1.0 0.0 0.0 0.0
0.0 1.0 0.0 0.0
0.0 0.0 1.0 0.0
0.0 0.0 0.0 1.0
1.0 0.0 0.0 0.0
0.0 1.0 0.0 0.0
0.0 0.0 1.0 0.0
0.0 0.0 0.0 1.0
输出
0.6698 0.3302 0.6698 0.3302
0.3302 0.6698 0.3302 0.6698
说明
dmodel=4,h=2,每头维度 dk=2。Wq=Wk=Wv=Wo=I(单位矩阵)。
拆分为 2 个头,每头处理 2 维。各头独立计算 Attention 后拼接,经输出投影得到最终结果。
四个投影矩阵按样例解释写成单位矩阵。
© CodeFun2000 · 使用条款
Scan the QR code below with WeChat to sign in
First-time scan will create your account automatically
请使用微信扫描下方二维码完成注册