flashdreams-postprocessing

द्वारा nvidia

FlashDreams वीडियो पोस्ट-प्रोसेसिंग प्रोसेसर, सत्र, प्रीसेट और रनर स्ट्रीम वायरिंग जोड़ें या संशोधित करें। नया VideoPostProcessorConfig लागू करते समय उपयोग करें /…

npx skills add https://github.com/nvidia/flashdreams --skill flashdreams-postprocessing

FlashDreams Post-Processing

Use this skill when adding a video post-processor or changing the runner post-processing stream. The reference implementation is integrations/flashvsr/flashvsr/postprocess.py.

Mental Model

A post-processor is usually three classes, not one class inheriting everything:

  • VideoPostProcessorConfig: serializable config and CLI surface. It sets _target to the processor factory and declares fields, output_spec(), requires_all_ranks(), and validate_execution().
  • VideoPostProcessor: lightweight factory created from config. Its job is start(spec) -> VideoPostProcessorSession.
  • VideoPostProcessorSession: mutable per-stream runtime. It owns buffers, caches, lazy model instances, counters, and process() / flush().

Keep stream state in the session. Do not store per-rollout mutable state on the config or processor factory.

Implementation Steps

  1. Pick a home:

    • Generic reusable post-processing belongs under flashdreams/flashdreams/infra/postprocess/.
    • Model-specific processors belong in their integration, for example integrations/<name>/<pkg>/postprocess.py.
  2. Define a config subclass:

    @dataclass(kw_only=True)
    class MyPostProcessorConfig(VideoPostProcessorConfig):
        _target: type["MyPostProcessor"] = field(
            default_factory=lambda: MyPostProcessor
        )
    
        scale: int = 2
    
        def output_spec(self, input_spec: VideoSpec) -> VideoSpec:
            return VideoSpec(
                height=input_spec.height * self.scale,
                width=input_spec.width * self.scale,
                fps=input_spec.fps,
                channels=input_spec.channels,
            )
    

    Override:

    • output_spec() when spatial size, channels, or timing changes.
    • requires_all_ranks() when the processor must run on nonzero ranks under torchrun.
    • validate_execution() to reject unsupported distributed or shape modes early.
  3. Define the processor factory:

    class MyPostProcessor(VideoPostProcessor[MyPostProcessorConfig]):
        def start(self, spec: VideoSpec) -> VideoPostProcessorSession:
            return _MyPostProcessorSession(self.config, spec)
    
  4. Define the session:

    class _MyPostProcessorSession(VideoPostProcessorSession):
        def __init__(self, config: MyPostProcessorConfig, spec: VideoSpec) -> None:
            self._config = config
            self._spec = spec
            self._buffer: Tensor | None = None
    
        def process(self, chunk: VideoChunk) -> list[VideoChunk]:
            ...
    
        def flush(self) -> list[VideoChunk]:
            ...
    

    process() is synchronous but may return []: that means it consumed the input chunk and is buffering frames until a later chunk or flush() can complete an output window.

  5. Handle layouts at the boundary:

    • Accept VideoChunk.tensor in chunk.layout.
    • Use to_bvtchw() only as a generic boundary helper.
    • Convert once into the processor's native layout, make it contiguous if the model kernels require that, and keep internal buffers in that native layout.
    • Document any forced .contiguous() because it can copy.
  6. Return VideoChunks:

    • Preserve [-1, 1] value range unless the API is intentionally changed.
    • Set the correct layout.
    • Carry metadata only if it helps downstream processors or provenance.
  7. Register presets when users should select it from CLI:

    [project.entry-points."flashdreams.postprocess_presets"]
    "my-postprocessor-v1" = "my_pkg.postprocess:POSTPROCESS_PRESET_MY_V1"
    

    The exported object must be a VideoPostProcessorConfig, for example:

    POSTPROCESS_PRESET_MY_V1 = MyPostProcessorConfig(...)
    

    Users select it with --postprocess.preset my-postprocessor-v1.

Runner Interaction

Runners create a VideoPostprocessStream through create_runner_postprocess_stream(). The stream:

  • creates one chain session for whole-stream processing, or one session per view when postprocess_per_view=True;
  • calls session.process(VideoChunk(...)) for each AR output;
  • turns [] into a zero-frame tensor so process() remains tensor-only;
  • skips collecting zero-time tensors in _append_if_nonempty();
  • calls flush() once at end-of-stream and appends any tail output.

Use postprocess_output_layout to describe the runner's decoded output layout. Use postprocess_per_view=True for bvtchw outputs when each camera/view needs an independent processor session.

Tests

Add CPU-safe tests unless the behavior genuinely requires a GPU:

  • Config/preset discovery: flashdreams/tests/test_postprocess_presets.py.
  • Stream contract and buffering: flashdreams/tests/test_postprocess_stream.py.
  • Processor-specific CPU fakes: integrations/<name>/tests/test_postprocess.py.
  • Runner distributed skip/all-rank behavior: flashdreams/tests/test_runner_postprocess.py.

Every pytest test must use exactly one marker: ci_cpu, ci_gpu, or manual. Prefer fake processor builders for CPU tests instead of loading checkpoints.

Useful focused validation:

uv run pytest flashdreams/tests/test_runner_postprocess.py \
  flashdreams/tests/test_postprocess_stream.py \
  flashdreams/tests/test_postprocess_presets.py \
  integrations/<name>/tests/test_postprocess.py

nvidia की और Skills

compileiq-debug
nvidia
उपयोग करें जब कुछ गलत हो: Search() हैंग हो जाता है, सभी मूल्यांकन INVALID_SCORE लौटाते हैं, स्कोर में सुधार नहीं हो रहा है, हर कॉन्फ़िगरेशन एक ही संख्या लौटाता है, ptxas त्रुटियाँ…
create-github-pr
nvidia
gh CLI का उपयोग करके GitHub पुल रिक्वेस्ट बनाएँ। जब उपयोगकर्ता नया PR बनाना चाहता है, कोड समीक्षा के लिए सबमिट करना चाहता है, या पुल रिक्वेस्ट खोलना चाहता है, तब उपयोग करें। ट्रिगर कीवर्ड -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
अन्य खुले मुद्दों को स्कैन करता है ताकि उन मुद्दों को ढूंढ सके जिन्हें कोई दिया गया PR ठीक कर सकता है या गलती से तोड़ सकता है। आसन्न-सुधार अवसरों और विरोधाभास जोखिमों को file:line… के साथ आउटपुट करता है।
fhir-basics
nvidia
एजेंटों को सिखाता है कि FHIR R4 APIs कैसे काम करते हैं, कौन से संसाधन उपलब्ध हैं, उन्हें खोज मापदंडों के साथ कैसे क्वेरी करें, और सभी प्रतिक्रिया प्रारूपों को सही ढंग से कैसे पार्स करें…
compileiq-validate-result
nvidia
खोज पूरी होने के बाद और किसी स्पीडअप का दावा करने या ACF भेजने से पहले उपयोग करें। dump_results CSV लोड करता है, शीर्ष-K उम्मीदवारों (एकल-उद्देश्य) को निकालता है…
changelog-audit
nvidia
रिलीज़ से पहले Warp CHANGELOG.md का ऑडिट करें: खोई हुई प्रविष्टियाँ पुनर्प्राप्त करें, उपयोगकर्ता प्रभाव के अनुसार क्रमबद्ध करें, प्रविष्टि भाषा को परिष्कृत करें, लाइन-रैप करें, और (रिलीज़-ब्रांच मोड) तुलना बढ़ाएँ…
maintain-dynamic-plugins
nvidia
NeMo Relay डायनामिक प्लगइन लोडर, मैनिफेस्ट, रस्ट नेटिव SDK, gRPC वर्कर प्रोटोकॉल, पायथन वर्कर SDK, दस्तावेज़, परीक्षण और रिलीज़ वर्कफ़्लो कवरेज बनाए रखें
dgx-diagnose
nvidia
सामान्य DGX Station GB300 समस्याओं का निदान करें — CUDA क्रैश, गलत-GPU लक्ष्यीकरण, vLLM/SGLang कंटेनर बग, MIG स्थिति समस्याएं, NVLink/Fabric Manager त्रुटियां,…