flashdreams-postprocessing

par nvidia

Ajouter ou modifier les processeurs de post-traitement vidéo FlashDreams, les sessions, les préréglages et le câblage des flux du runner. À utiliser lors de l'implémentation d'un nouveau VideoPostProcessorConfig /…

npx skills add https://github.com/nvidia/flashdreams --skill flashdreams-postprocessing

FlashDreams Post-Processing

Use this skill when adding a video post-processor or changing the runner post-processing stream. The reference implementation is integrations/flashvsr/flashvsr/postprocess.py.

Mental Model

A post-processor is usually three classes, not one class inheriting everything:

  • VideoPostProcessorConfig: serializable config and CLI surface. It sets _target to the processor factory and declares fields, output_spec(), requires_all_ranks(), and validate_execution().
  • VideoPostProcessor: lightweight factory created from config. Its job is start(spec) -> VideoPostProcessorSession.
  • VideoPostProcessorSession: mutable per-stream runtime. It owns buffers, caches, lazy model instances, counters, and process() / flush().

Keep stream state in the session. Do not store per-rollout mutable state on the config or processor factory.

Implementation Steps

  1. Pick a home:

    • Generic reusable post-processing belongs under flashdreams/flashdreams/infra/postprocess/.
    • Model-specific processors belong in their integration, for example integrations/<name>/<pkg>/postprocess.py.
  2. Define a config subclass:

    @dataclass(kw_only=True)
    class MyPostProcessorConfig(VideoPostProcessorConfig):
        _target: type["MyPostProcessor"] = field(
            default_factory=lambda: MyPostProcessor
        )
    
        scale: int = 2
    
        def output_spec(self, input_spec: VideoSpec) -> VideoSpec:
            return VideoSpec(
                height=input_spec.height * self.scale,
                width=input_spec.width * self.scale,
                fps=input_spec.fps,
                channels=input_spec.channels,
            )
    

    Override:

    • output_spec() when spatial size, channels, or timing changes.
    • requires_all_ranks() when the processor must run on nonzero ranks under torchrun.
    • validate_execution() to reject unsupported distributed or shape modes early.
  3. Define the processor factory:

    class MyPostProcessor(VideoPostProcessor[MyPostProcessorConfig]):
        def start(self, spec: VideoSpec) -> VideoPostProcessorSession:
            return _MyPostProcessorSession(self.config, spec)
    
  4. Define the session:

    class _MyPostProcessorSession(VideoPostProcessorSession):
        def __init__(self, config: MyPostProcessorConfig, spec: VideoSpec) -> None:
            self._config = config
            self._spec = spec
            self._buffer: Tensor | None = None
    
        def process(self, chunk: VideoChunk) -> list[VideoChunk]:
            ...
    
        def flush(self) -> list[VideoChunk]:
            ...
    

    process() is synchronous but may return []: that means it consumed the input chunk and is buffering frames until a later chunk or flush() can complete an output window.

  5. Handle layouts at the boundary:

    • Accept VideoChunk.tensor in chunk.layout.
    • Use to_bvtchw() only as a generic boundary helper.
    • Convert once into the processor's native layout, make it contiguous if the model kernels require that, and keep internal buffers in that native layout.
    • Document any forced .contiguous() because it can copy.
  6. Return VideoChunks:

    • Preserve [-1, 1] value range unless the API is intentionally changed.
    • Set the correct layout.
    • Carry metadata only if it helps downstream processors or provenance.
  7. Register presets when users should select it from CLI:

    [project.entry-points."flashdreams.postprocess_presets"]
    "my-postprocessor-v1" = "my_pkg.postprocess:POSTPROCESS_PRESET_MY_V1"
    

    The exported object must be a VideoPostProcessorConfig, for example:

    POSTPROCESS_PRESET_MY_V1 = MyPostProcessorConfig(...)
    

    Users select it with --postprocess.preset my-postprocessor-v1.

Runner Interaction

Runners create a VideoPostprocessStream through create_runner_postprocess_stream(). The stream:

  • creates one chain session for whole-stream processing, or one session per view when postprocess_per_view=True;
  • calls session.process(VideoChunk(...)) for each AR output;
  • turns [] into a zero-frame tensor so process() remains tensor-only;
  • skips collecting zero-time tensors in _append_if_nonempty();
  • calls flush() once at end-of-stream and appends any tail output.

Use postprocess_output_layout to describe the runner's decoded output layout. Use postprocess_per_view=True for bvtchw outputs when each camera/view needs an independent processor session.

Tests

Add CPU-safe tests unless the behavior genuinely requires a GPU:

  • Config/preset discovery: flashdreams/tests/test_postprocess_presets.py.
  • Stream contract and buffering: flashdreams/tests/test_postprocess_stream.py.
  • Processor-specific CPU fakes: integrations/<name>/tests/test_postprocess.py.
  • Runner distributed skip/all-rank behavior: flashdreams/tests/test_runner_postprocess.py.

Every pytest test must use exactly one marker: ci_cpu, ci_gpu, or manual. Prefer fake processor builders for CPU tests instead of loading checkpoints.

Useful focused validation:

uv run pytest flashdreams/tests/test_runner_postprocess.py \
  flashdreams/tests/test_postprocess_stream.py \
  flashdreams/tests/test_postprocess_presets.py \
  integrations/<name>/tests/test_postprocess.py

Plus de skills de nvidia

compileiq-debug
nvidia
Utilisez quand quelque chose ne va pas : Search() bloque, toutes les évaluations retournent INVALID_SCORE, les scores ne s'améliorent pas, chaque configuration retourne le même nombre, erreurs ptxas…
create-github-pr
nvidia
Créer des pull requests GitHub en utilisant l'interface en ligne de commande gh. Utiliser lorsque l'utilisateur souhaite créer une nouvelle PR, soumettre du code pour révision, ou ouvrir une pull request. Mots-clés de déclenchement -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Analyse les autres problèmes ouverts pour trouver ceux qu’une PR donnée pourrait également corriger ou casser accidentellement. Génère des opportunités de correctifs adjacents et des risques de contradiction avec fichier:ligne…
fhir-basics
nvidia
Apprend aux agents comment fonctionnent les API FHIR R4, quelles ressources sont disponibles, comment les interroger avec des paramètres de recherche, et comment analyser correctement tous les formats de réponse…
compileiq-validate-result
nvidia
Utiliser APRÈS qu'une recherche soit terminée et AVANT de réclamer un accélérateur ou d'expédier un ACF. Charge le CSV dump_results, extrait les K meilleurs candidats (mono-objectif)…
changelog-audit
nvidia
Auditer le CHANGELOG.md de Warp avant une publication : récupérer les entrées perdues, trier par impact utilisateur, affiner le langage des entrées, ajuster les retours à la ligne et (en mode branche de publication) mettre à jour la comparaison…
maintain-dynamic-plugins
nvidia
Maintenir les chargeurs de plugins dynamiques NeMo Relay, les manifestes, les SDK natifs Rust, le protocole worker gRPC, le SDK worker Python, la documentation, les tests et la couverture du workflow de publication
dgx-diagnose
nvidia
Diagnostiquer les problèmes courants du DGX Station GB300 — plantages CUDA, ciblage incorrect du GPU, bugs de conteneur vLLM/SGLang, problèmes d'état MIG, erreurs NVLink/Fabric Manager,…