ByteAction: Byte-space Action Recognition Foundation Model
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608.21837v1 Announce Type: new Abstract: Bitstream-corrupted Harsh Visual Understanding (BcHVU) aims to understand harshly degraded videos originally decoded from a severely corrupted bitstrea...
arXiv:2610.08414v1 Announce Type: new Abstract: Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded...
arXiv:2608.29212v1 Announce Type: cross Abstract: Existing video watermarking systems are symmetric: the party that can verify a mark holds the extractor weights or generator secret and can therefore...
arXiv:2609.39623v1 Announce Type: new Abstract: The proliferation of high-fidelity generative editing models has made it possible to inject violent or sexual content into otherwise ordinary images wh...
COVER is a new video watermarking method that targets codec compression as its primary design goal. It embeds the watermark payload in the latent space of a frozen generative video autoencoder and recovers it by re‑encoding the received video into the same latent space. Using a differentiable codec surrogate bank, COVER achieves high bit accuracy across multiple codecs while keeping marked videos visually close to the originals.
arXiv:2410. 19553v2 Announce Type: replace-cross Abstract: This paper explores the impact of occlusions in video action detection.