The output size of a convolutional layer is calculated using the formula: output = ⌊(input − kernel + 2 × padding) / stride⌋ + 1. For example, a 224×224 input image with a 3×3 kernel, stride 1, and padding 1 produces a 224×224 output (same padding). With stride 2 and no padding, the same kernel produces a 111×111 output, halving the spatial dimensions. The number of output channels equals the number of filters (kernels) in the layer. For pooling layers, the same formula applies using the pool size as the kernel. Understanding this formula is essential for designing CNN architectures — incorrect dimension calculations cause shape mismatch errors that are among the most common bugs in deep learning code. Transposed convolutions (deconvolutions) use a related formula for upsampling: output = (input − 1) × stride − 2 × padding + kernel.
Built with care by Alpiaal