Skip to content

InternVideo-Next Multi-modality probes #318

Description

@xuridi01

Hello @Revliter ,
Thank you so much for your previous response on issue #312 — I really appreciate the time you've taken to help. I've been digging deeper into the text encoder training for InternVideo-Next and have run into a few questions I'd love your input on.

Question 1 — Correct Training Setup: Paper vs. InternVideo2
In the paper, you mention freezing the ViT backbone and training only the text encoder. However, in issue #312 you pointed me toward the InternVideo2 multi-modality training, which uses a slightly different setup:

  • Vision backbone → fully frozen
  • Text backbone → fully frozen
  • clip-projector (vision side) → unfrozen
  • Alignment layer added on the vision side

Could you clarify which approach is correct for reproducing InternVideo-Next zero shot t2v results? Specifically: Should I follow the InternVideo2 setup exactly, or Adapt it to better match the paper — e.g., unfreeze the text backbone, and optionally freeze/unfreeze the clip-projector and add alignment on the text and/or vision side?

Question 2 — Dimension Alignment with SigLIP2 1B Teacher
You mentioned that SigLIP2 1B (giant opt) was used as a teacher in Stage 1 pretraining. However, its embedding dimensionality is quite different from the resulting InternVideo-Next vision encoder. How was dimension alignment handled between the two models?
Additionally — and I'm not sure if you tried this — i would expect the InternVideo-Next vision encoder shift away from SigLIP2's embedding space after Stage 2, making the two spaces incomparable at that point right?

Question 3 — Text-Side Training Settings, Epochs, and Room for Improvement
A few related sub-questions here:

  • Training config: Do the text-side training settings (temperature, epochs, weight decay, learning rate) fully follow the InternVideo2 configs?
  • Epoch discrepancy: In the paper, zero-shot T2V results are compared against InternVideo2 CLIP-L/14, which was trained for 3 epochs, whereas the InternVideo-Next multi-modality probe (Section: Multi-modality Tasks) mentions 5 epochs. Could you clarify this difference?
  • Are these results final? You refer to these experiments as probes — do you believe there is room for improvement with further tuning (e.g., dataset size, text encoder size, hyperparameters), or are the reported numbers the expected ceiling for this configuration?

I find this work incredibly insightful and plan to use the vision encoder in my diploma thesis given its strong potential. These clarifications would really help me move forward.
Thank you so much in advance for your time and help!

Activity

  1. bejvisek commented on Apr 29, 2026

    @bejvisek

    Wow, it's a great set of questions, some of which I was asking myself when working with the repo!

    It would be super helpful to clarify this and it would help me a lot to replicate the great work you did @Revliter

    Thank you!

  2. Dajvid commented on May 3, 2026

    @Dajvid

    +1 from me, great work @Revliter! I’d love to get your insights on this.

  3. getmoments-com commented on May 5, 2026

    @getmoments-com

    Hi @Revliter, I also vote for this :-) It's a great work and we would love to be able to repeat it as in the paper!

  4. Revliter commented on May 6, 2026

    @Revliter
    Collaborator

    Hi,

    Thanks a lot for your thoughtful questions and for your interest in our work — really glad to hear you're exploring InternVideo-Next so deeply.

    Q1 — Training setup

    You are right that the description in the paper can be a bit unclear due to space limitations. In practice, the setup is consistent with InternVideo2-L: we freeze both the vision and text backbones, and only train the projector and the alignment layer.

    The main reason for keeping the text encoder frozen is that the current video-text data we use is still relatively noisy / limited in quality. Fully unfreezing the text backbone tends to degrade the original text representation ability inherited from MobileCLIP. By freezing it, we preserve its strong text embedding space while aligning the vision features to it.

    That said, we do believe this is not a fundamental limitation — with higher-quality video-text data or better hyperparameter tuning, it could be beneficial to partially or fully unfreeze the text encoder.

    So for reproducing the zero-shot T2V results of InternVideo-Next, it would be best to follow this setup.

    Q2 — Dimension alignment with SigLIP2

    We use MLP decoder to project features into a higher-dimensional space for alignment with the SigLIP2 1B teacher.

    Regarding your intuition — yes, that is correct. After Stage 2, the embedding space shifts, especially in the motion domain. In fact, this can be indirectly observed from the performance gap: most image-based encoders perform poorly on Something-Something V2, while our Stage 2 model performs much better. This suggests that the representation has moved away from the original image-aligned embedding space. We believe that the model learns about 'captioning each frame separately' in Stage 1 and 'understand the full video dynamics' in Stage 2.

    Q3 — Text-side training details and potential improvements

    The text-side training setup is not exactly the same as InternVideo2. Since InternVideo-Next does not see large-scale video-text alignment during training, we found it helpful to train for more epochs to better probe its text-retrieval capability after certain warm-up. Some hyperparameters (e.g., temperature, learning rate, etc.) were also adjusted accordingly, as the embedding space has different properties under this text-free training paradigm.

    We have also experimented with using LLM-based embeddings (Qwen3-embedding) in this LiT-style setup, and observed similar performance. The model's retrieval ability may still be bounded by the SigLip2's ability in Stage 1. One possible direction for improvement is to incorporate some amount of high-quality video-text data in Stage 1 pre-training. That said, the main goal of this work is to explore the upper bound of video-only (text-free) training, rather than to completely replace video-text training. We are optimistic that future work can better combine this paradigm with video-text alignment for further gains. (maybe not with the stage-2 embedding, but with the stage-1 training.)

    Additional clarification

    One thing we would like to clarify is that we did not initially expect such a strong focus from the community on the text-side performance. This work is primarily intended as an exploration of video-only (text-free) pre-training and its effectiveness on general video understanding tasks, especially those that do not rely on text supervision but are more video-centric or related to dynamics more, such as action/motion recognition, depth estimation, tracking, ... As a result, we did not place as much emphasis on optimizing text-related retrieval performance in this study originally. There might be space towards better video-text performance with more advanced methods designed for video-text coherence.

    Hope this helps, and best of luck with your thesis. We will release the pre-training code before 5.20.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions