← Back to all news
arXiv Computer Vision September 22, 2026 By Mena Kamel, Natalie Won, Amrut Sarangi, Sven Jager, Albert Pla Planas

STA-TFM: Spatio-Temporal Aggregation Across Views TransForMer for Pose Estimation

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • computer-vision

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
Jul 13

HandFlow: Fully Generative 4D Hand Recovery with Flow Matching

Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.

llmsdiffusionbenchmarks
More like this →
arXiv AI
Jul 14

HandFlow: Fully Generative 4D Hand Recovery with Flow Matching

arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.

By Mingxi Xu, Bowen Duan, Yi Gu, Zhengyang Shen, Renjing Xu, Yutao Yue
llmsdiffusionbenchmarks
More like this →
arXiv Computer Vision
Sep 22

Estimating Accurate Hand Pose in Camera Space with Vision Transformer

arXiv:2609.24424v1 Announce Type: new Abstract: Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict...

By Kaiwen Ren, Yiran Jiang, Yongjing Ye, Shihong Xia
llmsragcomputer-visionbenchmarks
More like this →
arXiv Computer Vision
Sep 15

AG-EgoPose: Spatially Anchored Residual Correction with Action Context for Monocular Egocentric 3D Pose Estimation

arXiv:2603.25175v2 Announce Type: replace Abstract: Monocular egocentric 3D pose estimation is difficult because severe foreshortening, self-occlusion, and a restricted field of view often remove the...

By Md Mushfiqur Azam, John Quarles, Kevin Desai
llmscomputer-visionfine-tuning
More like this →
arXiv Computer Vision
1d ago

HARMONI: Aligning Human and Scene Priors for Multi-View 4D Reconstruction

arXiv:2603.12789v3 Announce Type: replace Abstract: Recent advances in 3D foundation models have enabled joint reconstruction of humans and their surrounding environments. However, combining independ...

By Sangmin Kim, Minhyuk Hwang, Geonho Cha, Dongyoon Wee, Jaesik Park
computer-vision
More like this →
arXiv Computer Vision
Sep 2

On the Role of Rotation Equivariance in Monocular 2D-to-3D Human Pose Lifting

arXiv:2601.13913v3 Announce Type: replace Abstract: We consider monocular 3D human pose estimation (HPE), where the goal is to predict 3D human skeletal joints from a single 2D image, typically via 2...

By Pavlo Melnyk, Cuong Le, Urs Waldmann, Per-Erik Forss\'en, Bastian Wandt
computer-visionbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea