The core idea of the Transformer architecture is to completely abandon traditional recurrent neural networks and convolutional neural networks, relying solely on the self-attention mechanism to capture global dependencies within an input sequence, thereby achieving an extremely high degree of parallelization.