Pruna's P-Video-2-Pro Spreads MiniMax H3 Across 9 Inference Platforms
Pruna's P-Video-2-Pro launched across 9 inference platforms at once, betting on distribution over chasing the fastest raw speed record.
Pruna AI released P-Video-2-Pro, a MiniMax H3-based video model, live today across nine separate inference partners simultaneously: Replicate, Cloudflare, Each Labs, Scenario, Tellers AI, WaveSpeed AI, Wiro AI, Runware, and Together AI. The speed numbers are genuinely solid, but they're not the record-setter among H3 variants covered here. What actually sets this release apart is how many places you can access it from at launch.
What P-Video-2-Pro Actually Does
Today, we launch P-Video-2-Pro, based on MiniMax H3: The best quality. The best speed.
β Pruna AI (@PrunaAI) September 17, 2026
Text or image in. Video and audio out.
β’ ππ»π½ππ: Text or first-frame image, with optional last-frame conditioning
β’ π¦π½π²π²π±: ~2.0s for a 5s clip at 480p Β· ~4.3s at 768p in Speed modeβ¦ pic.twitter.com/M184RswHHW
The model accepts text or a first-frame image as input, with optional last-frame conditioning for controlling how a clip ends, and outputs video with synced audio. Two modes trade speed for quality directly: Speed mode generates a 5-second clip at 480p in roughly 2.0 seconds, or at 768p in about 4.3 seconds. Quality mode presumably takes longer in exchange for better output, though Pruna hasn't published those specific numbers yet.
Pricing runs $0.02 per second in Speed mode and $0.04 per second in Quality mode, straightforward, transparent per-second billing rather than credit-based tiers. Additional controls include seed selection, aspect ratio, and three levels of prompt upsampling (Off, Turbo, or Max) for automatically expanding a short prompt into more detailed generation instructions.
The Real Differentiator Is Distribution, Not Speed

Compare P-Video-2-Pro's headline number, 2.0 seconds for a 5-second 480p clip, against NVIDIA's Sol-H3, which generates a 5-second 768p clip with audio in 1.653 seconds on 8x B300 hardware. Sol-H3 still holds the fastest verified result among the H3 speed projects covered here. P-Video-2-Pro isn't claiming to beat that.
What it does claim is breadth: nine inference partners live simultaneously at launch, plus a free public Playground for testing before committing to API access. That's a meaningfully different strategy than most of the H3 acceleration projects so far, which have generally launched through a single hosting platform or as self-hosted open weights requiring your own infrastructure.
Who Pruna Actually Is

Pruna AI describes itself as a model laboratory and inference provider, building its own "performance models" tuned for the balance of speed, cost, and quality, then distributing them across a wide network of partner platforms rather than picking one exclusive host. P-Video-2-Pro isn't Pruna's first release in this line either; the company has been iterating on P-Video generally, including workflows for chaining multiple generated segments into longer-form video using an LLM to plan scenes and the last frame of each clip as the input for the next.
Pruna states the model's benchmarks are being validated by Datapoint AI, Rapidata AI, and Design Arena, with results to be published soon, worth checking once those land to see how the quality claims hold up against independent testing rather than the company's own numbers alone.
Competitive Context

This is the seventh distinct approach to MiniMax H3 speed and accessibility covered here, joining fal's H3 Max, NVIDIA's Sol-H3, FastH3, VDN, Alibaba's PDD/Acc-LoRAs, and LightX2V.
Where those projects mostly compete on raw technical approach, distillation, sparse attention, architectural rebuilds, Pruna's contribution is closer to infrastructure and reach: getting a well-optimized version of H3 in front of developers already building on whichever platform they've already chosen, rather than requiring everyone to adopt one specific hosting service.
The Signal in the Noise
Nine simultaneous launch partners is a real signal about how commoditized fast H3 inference is becoming. When a company can ship the same optimized model across nearly a dozen platforms on day one, speed alone is no longer the differentiator it was a month ago; distribution and ease of integration are becoming the new competitive layer on top of it.
Given how many MiniMax H3 speed and access options now exist, does broad multi-platform availability matter more to you than squeezing out the last fraction of a second in raw generation speed?
The Details
- Release: P-Video-2-Pro, by Pruna AI, based on MiniMax H3
- Input: text or first-frame image, with optional last-frame conditioning
- Speed: ~2.0s for a 5s clip at 480p, ~4.3s at 768p, in Speed mode
- Pricing: $0.02/second (Speed mode), $0.04/second (Quality mode)
- Controls: first/last frame, seed, aspect ratio, prompt upsampling (Off/Turbo/Max)
- Launch inference partners (9): Replicate, Cloudflare, Each Labs, Scenario, Tellers AI, WaveSpeed AI, Wiro AI, Runware, Together AI
- Benchmark validation (results pending): Datapoint AI, Rapidata AI, Design Arena
- Access: free Playground, plus API via Pruna's dashboard or any listed partner platform