Posts

Showing posts with the label #Transformer

Cross Attention in Decoder Block of Transformer

Image
  Notice where the cross attention is marked, 2 arrows are coming from encoder block, and one is coming from decoder block. Why do we need to consider Encoder Block? Now, lets say we have predicted 2 words, and we need to predict the 3rd word? It will depend on what? Ofc, first 2 words of decoder block, and original sentence context from the Encoder block. So, we need to figure out the relationship between these two. How will we get the relationship? q : Hindi (from Decoder Block) k : Eng (from Encoder Block) v : Eng (from Encoder Block)

Encoder Architechture - Transformer

Image
 Transformer Architecture : Transformer Architecture has 2 parts : 1. Encoder 2. Decoder Each encoder and docoder part repeats 6 time. (mentioned in Attention is all you need paper) They found that this is the best fit value from experiments. Each encoder block is identical to others, same for decoder. Now, lets zoom in to the Encoder Block. 6 Encoder Blocks: Input would look like this: Input : sentence Output : positional Encoded Vector Now, lets focus on Multihead attention and Normalization part: Input : Output of first step (positional encoded vector) Output :  Normalized vectors Why is there residual connection and why do we need to add input and final vector after multihead attention? Answer is not mentioned in Attention is all you need paper and nobody has an idea on the internet.          People assume 2 possibilities: 1. For stable training 2. If the current layer messes up, we still have some good original data Now, let's focus on Feed-Forwar...