esvit icon indicating copy to clipboard operation
esvit copied to clipboard

Allow arbitrary image sizes and upstream changes from Swin-Transformer-Object-Detection

Open vadimkantorov opened this issue 4 years ago • 1 comments

It is useful in object detection context to allow arbitrary sizes by doing dynamic mask computation (probably possible only with relative position encoding).

These kinds of edits were done in https://github.com/SwinTransformer/Swin-Transformer-Object-Detection and in https://github.com/megvii-research/SOLQ/. It would be nice if you upstreamed these changes. This will simplify trying out ESviT checkpoints as pretraining for object detection.

Also, fyi I created a similar issue in SimMIM: https://github.com/microsoft/SimMIM/issues/13. Overall, having some stable version of swin_transformer.py somewhere (maybe even in main SwinTransformer/Swin-Transformer repo?) supporting dynamic masking would help a lot :)

Thanks!

vadimkantorov avatar Jan 26 '22 10:01 vadimkantorov

Hi,do you have ckpt and train logs , can you share with me ? I got an error ,when I download them.

sym0926 avatar Aug 14 '24 11:08 sym0926