We handle the elemental incompatibility of attention-based encoder-decoder (AED) fashions with long-form acoustic encodings. AED fashions skilled on segmented utterances study to encode absolute body positions by exploiting restricted acoustic context past section boundaries, however fail to generalize when decoding long-form segments the place these cues vanish. The mannequin loses capability to order acoustic encodings as a result of permutation invariance of keys and values in cross-attention. We suggest 4 modifications: (1) injecting express absolute positional encodings into cross-attention for every decoded section, (2) long-form coaching with prolonged acoustic context to eradicate implicit absolute place encoding, (3) section concatenation to cowl various segmentations wanted throughout coaching, and (4) semantic segmentation to align AED-decoded segments with coaching segments. We present these modifications shut the accuracy hole between steady and segmented acoustic encodings, enabling auto-regressive use of the eye decoder.

