AReA: Adaptive Receptive-Field Attention for Efficient High-Resolution Diffusion Transformers

University of TorontoVector Institute

* Equal contribution

Previous workSEGA

scroll
High-resolution generation. AReA enables training-free and efficient high-resolution generation across FLUX.1, Qwen-Image, FLUX.2, and Ideogram 4. Move the pointer over the image to magnify it · on a phone, drag the lens
High-Resolution Image Generation
up to 6× the native side, 67 megapixels
Efficient
fewer keys per query, up to 3.2× faster
Training-free
pretrained weights untouched
RoPE unchanged
no RoPE extrapolation, applied without modification

Abstract

Generating images beyond the training resolution of Diffusion Transformers (DiTs) remains challenging, as the number of tokens grows rapidly with resolution, full attention becomes expensive, and outputs often suffer from structural distortions and degraded detail. We identify uniform full attention as the common source of both problems. Attending to all tokens at every step causes quadratic cost and exposes the model to more tokens and longer relative distances than seen during training, pushing it out of distribution. We observe that attention’s effective receptive field narrows during denoising, from global in early steps that form structure to local in later steps that refine details. We propose AReA, a training-free attention mechanism whose receptive field evolves from coarse to fine, following the denoising progression. In early steps, AReA applies cross-scale coarse attention over downsampled keys and values, capturing global context, while in later steps it applies local attention within progressively shrinking windows. Without additional parameters or training, experiments show that AReA consistently outperforms state-of-the-art training-free baselines across target resolutions and different models, while also being faster and more efficient.

In motion

One image through two phases

method video · coming soon
A real FLUX.1 run at 4096², step by step: the problem, the observation, then AReA’s two phases.

The problem

Full attention, far past training

Every token attends to every token at every step: the cost grows quadratically, and the model sees more tokens and longer relative distances than in training, which pushes it out of distribution.

A tall model in a quilted saffron overcoat poses on a broad concrete stairway in directional morning sun, with an elegant silhouette separated clearly from a restrained background…

layout collapse
Full attention
fragmentation
UltraImage
incomplete subject
SigMa
distorted scales
SEGA
coherent structure, fine detail
AReA (ours)

Failure modes at high resolution. FLUX.1 at 4096 × 4096. Full attention keeps the model’s intended subject but collapses the layout, while RoPE-based methods keep a layout but fragment, lose, or distort the subject. AReA preserves both, and renders fine details.

The observation

Attention narrows

from global in early steps that form structure to local in later steps that refine details. So the receptive field that attention needs varies across denoising steps.

Distance to dominant keys
5.5×closer
14.62.7 tokens
late in denoising, the most important keys lie close to the query
Distance over all keys
1.4×closer
17.913.1 tokens

The idea

AReA: from coarse to fine

An attention receptive field that follows the denoising progression: global through pooled keys while the layout forms, then local windows that shrink as fine details emerge. AReA never computes full attention at the target resolution.

PHASE I · σ > σ*
Cross-scale coarse attention

keys and values pooled to one token per f × f block: every token still sees the whole image

PHASE II · σ < σ*
Adaptive windowed fine attention

local windows shrink as fine details emerge; windows over smooth regions split

global, coarse local, fine σ = 1 σ* 0

denoising · drag the marker

Pipeline of AReA. Denoising moves from coarse to fine across σ*. In Phase I, full-resolution queries attend to keys and values pooled to one token per f × f block on a native-resolution grid with a ramped f. In Phase II, attention runs in shrinking local windows, and windows over smooth regions of the Discrete Wavelet Transform (DWT) map are split to cut cost.

Phase I · coarse

Global, early

The early steps decide the global layout, so every token needs a receptive field that covers the whole grid. Queries stay at full resolution; keys and values are pooled.

pooled keys

One token per f × f block

Each block of the grid is represented by its centre token, for keys and values alike. Each of the N queries attends to all T + N/f² keys, at 1/f² of the cost, and the output stays at full resolution.

positions

On a native-resolution grid

The pooled keys sit on a coarse (W/f) × (H/f) grid and query coordinates are scaled by 1/f to match, so the relative distances RoPE encodes stay close to the model’s training range.

pooling ramp

Finer pooling as the layout forms

A large f keeps the pooled grid within the model’s native range but is too coarse for object boundaries; a small f duplicates subjects. AReA starts from the f0 that matches the native resolution and decreases it linearly over Phase I down to 2, each value snapped to a divisor of the grid, so the pooled keys get finer as the image’s own detail grows.

f̃t = f0 + (t − 1)/(Sc − 1) · (fmin − f0)

Phase II · fine

Local, late

Once the layout is set at σ*, fine details are local: each token interacts mainly with its nearby tokens. Windows start large for coherence and shrink every step, down to the model’s native range.

shrinking windows

Smaller every step

From about 70% of the grid’s longer side down to the native size, each step checked against a cost budget so no windowed step costs more than one full-attention pass.

wt = wstart · (wmin/wstart)(t−1)/(Sd−1)
DWT splits

Smooth windows split in four

A wavelet map of the current state measures each region’s detail. Windows over smooth regions split into quadrants; windows over detail, such as a face, keep their size.

Halo

A ring of neighbours

Each window’s keys and values are extended with a narrow ring of neighbouring tokens, so edge tokens see across the boundary. Window boundaries also shift randomly across steps, so seams never settle.

Global-KV Bridge

Distinctive tokens, shared

Tokens far from the global mean key, spread over the whole canvas, join every window’s keys in the same softmax, a fixed fraction of the window, so small objects are not duplicated.

Results

Coherent structure, fine detail

Direct-inference baselines stretch bodies, duplicate subjects and garble text at high resolution. AReA keeps the intended global structure while generating fine details, without altering positional encodings.

Efficiency

Faster as resolution grows

Every windowed step is cheaper than one full-attention pass, so the speedup grows with resolution, where attention cost increasingly dominates.

3.2×

faster than full attention at 6144² on FLUX.1

2.9×

faster at 6144² on Qwen-Image

Citation

BibTeX will appear here once the paper is public.