One token per f × f block
Each block of the grid is represented by its centre token, for keys and values alike. Each of the N queries attends to all T + N/f² keys, at 1/f² of the cost, and the output stays at full resolution.
Abstract
Generating images beyond the training resolution of Diffusion Transformers (DiTs) remains challenging, as the number of tokens grows rapidly with resolution, full attention becomes expensive, and outputs often suffer from structural distortions and degraded detail. We identify uniform full attention as the common source of both problems. Attending to all tokens at every step causes quadratic cost and exposes the model to more tokens and longer relative distances than seen during training, pushing it out of distribution. We observe that attention’s effective receptive field narrows during denoising, from global in early steps that form structure to local in later steps that refine details. We propose AReA, a training-free attention mechanism whose receptive field evolves from coarse to fine, following the denoising progression. In early steps, AReA applies cross-scale coarse attention over downsampled keys and values, capturing global context, while in later steps it applies local attention within progressively shrinking windows. Without additional parameters or training, experiments show that AReA consistently outperforms state-of-the-art training-free baselines across target resolutions and different models, while also being faster and more efficient.
In motion
The problem
Every token attends to every token at every step: the cost grows quadratically, and the model sees more tokens and longer relative distances than in training, which pushes it out of distribution.
A tall model in a quilted saffron overcoat poses on a broad concrete stairway in directional morning sun, with an elegant silhouette separated clearly from a restrained background…
Failure modes at high resolution. FLUX.1 at 4096 × 4096. Full attention keeps the model’s intended subject but collapses the layout, while RoPE-based methods keep a layout but fragment, lose, or distort the subject. AReA preserves both, and renders fine details.
The observation
from global in early steps that form structure to local in later steps that refine details. So the receptive field that attention needs varies across denoising steps.
The idea
An attention receptive field that follows the denoising progression: global through pooled keys while the layout forms, then local windows that shrink as fine details emerge. AReA never computes full attention at the target resolution.
keys and values pooled to one token per f × f block: every token still sees the whole image
local windows shrink as fine details emerge; windows over smooth regions split
denoising · drag the marker
Phase I · coarse
The early steps decide the global layout, so every token needs a receptive field that covers the whole grid. Queries stay at full resolution; keys and values are pooled.
Each block of the grid is represented by its centre token, for keys and values alike. Each of the N queries attends to all T + N/f² keys, at 1/f² of the cost, and the output stays at full resolution.
The pooled keys sit on a coarse (W/f) × (H/f) grid and query coordinates are scaled by 1/f to match, so the relative distances RoPE encodes stay close to the model’s training range.
A large f keeps the pooled grid within the model’s native range but is too coarse for object boundaries; a small f duplicates subjects. AReA starts from the f0 that matches the native resolution and decreases it linearly over Phase I down to 2, each value snapped to a divisor of the grid, so the pooled keys get finer as the image’s own detail grows.
f̃t = f0 + (t − 1)/(Sc − 1) · (fmin − f0)Phase II · fine
Once the layout is set at σ*, fine details are local: each token interacts mainly with its nearby tokens. Windows start large for coherence and shrink every step, down to the model’s native range.
From about 70% of the grid’s longer side down to the native size, each step checked against a cost budget so no windowed step costs more than one full-attention pass.
wt = wstart · (wmin/wstart)(t−1)/(Sd−1)A wavelet map of the current state measures each region’s detail. Windows over smooth regions split into quadrants; windows over detail, such as a face, keep their size.
Each window’s keys and values are extended with a narrow ring of neighbouring tokens, so edge tokens see across the boundary. Window boundaries also shift randomly across steps, so seams never settle.
Tokens far from the global mean key, spread over the whole canvas, join every window’s keys in the same softmax, a fixed fraction of the window, so small objects are not duplicated.
Results
Direct-inference baselines stretch bodies, duplicate subjects and garble text at high resolution. AReA keeps the intended global structure while generating fine details, without altering positional encodings.
Efficiency
Every windowed step is cheaper than one full-attention pass, so the speedup grows with resolution, where attention cost increasingly dominates.
faster than full attention at 6144² on FLUX.1
faster at 6144² on Qwen-Image
Citation
BibTeX will appear here once the paper is public.