Mind Video

Mind Video reconstructs high-quality video from brain activity (fMRI) via masked brain modeling and an augmented Stable Diffusion model.

Last verified:

Visit Mind Video

What is Mind Video?

Mind Video (MinD-Video) is an AI-powered framework that reconstructs high-quality videos directly from human brain activity captured through continuous fMRI data. Developed by researchers at the National University of Singapore and The Chinese University of Hong Kong, Mind Video bridges static image brain decoding and full video reconstruction from brain signals using a two-module pipeline: an fMRI encoder that learns spatiotemporal brain features through masked brain modeling and multimodal contrastive learning with spatiotemporal attention, and an augmented Stable Diffusion model with network temporal inflation to generate coherent video under fMRI guidance.

Mind Video achieved 85% accuracy in semantic classification and 0.19 SSIM, outperforming previous state-of-the-art methods by 45%. It reconstructs objects, animals, persons, motions, and scene dynamics, accepts arbitrary frame rates, and works with longer recordings to reconstruct longer videos. Accepted at NeurIPS 2023 (oral), it is designed for neuroscience and cognitive science researchers studying human visual perception and brain decoding.

Mind Video pricing

Pricing model: Free

MinD-Video is an open-source research framework available for free on GitHub. The codebase is publicly accessible under the jqin4749/MindVideo repository. Pre-trained checkpoints and preprocessed test data are available for download via Google Drive. No commercial pricing or paid tiers exist as this is academic research software. Users need their own GPU hardware (RTX3090 or higher) and must set up the conda environment to run the model locally.

Mind Video pros

  • 85% semantic accuracy outperforming previous SOTA by 45%
  • Two-module pipeline with fMRI encoder and augmented Stable Diffusion
  • Reconstructs arbitrary frame rates from 3 FPS to full 30 FPS
  • Handles diverse objects: animals, persons, basic objects, scene types
  • Accurately reconstructs motions like running, dancing, singing
  • Reconstructs scene dynamics including close-ups and fast-motion
  • Biologically plausible and interpretable model
  • Reflects established physiological processes in brain
  • Accepted for Oral Presentation at NeurIPS 2023
  • Uses publicly available HCP pre-training dataset
  • Unsupervised learning techniques for semantic space assimilation
  • Unique spatiotemporal attention mechanism for time simulation
  • Adversarial guidance for high-quality video generation
  • Can work with longer brain recordings for longer videos
  • High-definition video output at 256x256 resolution

Mind Video cons

  • Requires continuous fMRI data from expensive brain imaging equipment
  • Currently limited to 2 seconds at 3 FPS due to GPU memory
  • Needs RTX3090 or higher GPU for generation
  • 256x256 resolution limitation without more GPU memory
  • Complex two-module pipeline requires specialized knowledge
  • Only works with fMRI, not other brain recording methods
  • Public dataset limitation (Wen 2018 dataset only)
  • Research tool not yet available as public web service

Frequently asked questions about Mind Video

What is MinD-Video?

MinD-Video is a framework for high-quality video reconstruction from brain recording using continuous fMRI data. It uses a two-module pipeline with an fMRI encoder and augmented Stable Diffusion model to reconstruct cinematic videos that people are viewing from their brain activity.

How accurate is MinD-Video?

MinD-Video achieved 85% average accuracy in semantic classification tasks and 0.19 structural similarity index (SSIM), outperforming the previous state-of-the-art by 45%. This represents a significant breakthrough in brain decoding technology.

What equipment do I need to use MinD-Video?

To use MinD-Video, you need continuous fMRI data from brain recordings and a GPU with at least RTX3090 capabilities. The current samples are generated with one RTX3090, though more GPU memory enables longer videos at higher resolution and frame rates.

What types of videos can MinD-Video reconstruct?

MinD-Video can reconstruct various objects, animals, persons, and scene types. It accurately reconstructs motions like running, dancing, and singing, plus scene dynamics including close-ups of people, fast-motion scenes, and long-shot city views.

What is the two-module pipeline?

The two-module pipeline consists of an fMRI encoder that progressively learns brain features through masked brain modeling and multimodal contrastive learning with spatiotemporal attention, and an augmented Stable Diffusion model fine-tuned for video generation under fMRI guidance with network temporal inflation.

Can MinD-Video reconstruct longer videos?

Yes, MinD-Video can work with longer brain recordings to reconstruct longer videos with full frame rate (30 FPS) and higher resolution, though this requires more GPU memory. Current samples are limited to 2 seconds at 3 FPS due to GPU memory constraints.

Is MinD-Video available as a web service?

No, MinD-Video is an open-source research framework available on GitHub. Users must download the codebase, set up the conda environment, download checkpoints from Google Drive, and run it locally on their own GPU hardware.

What datasets does MinD-Video use?

MinD-Video uses the large-scale pre-training dataset from HCP (Human Connectome Project) and the target dataset Wen (2018) which contains videos and corresponding fMRI brain recordings from test subjects who watched the videos.

Is MinD-Video biologically plausible?

Yes, MinD-Video is shown to be biologically plausible and interpretable, reflecting established physiological processes in the brain. The model's design incorporates mechanisms that align with how the cerebral cortex processes spatiotemporal visual information.

What resolution and frame rate does MinD-Video produce?

Current samples are generated at 256x256 resolution at 3 FPS (2 seconds duration) using one RTX3090. However, the method can support arbitrary frame rates including full 30 FPS and higher resolutions if more GPU memory is available.

Categories

Use cases

Browse all AI tools on NeedAnAI