Neurise EN / Blog / Grok Imagine Video 1.5

Grok Imagine Video 1.5 from xAI: native audio and a price worth checking yourself

The model turns a still image into video with its own soundtrack. We set xAI's price list against OpenAI's and check what the "86 per cent cheaper than Sora" headline actually compares.

Grok Imagine Video 1.5 is an xAI model that turns a single still image into a video a few seconds long, with a soundtrack generated in the same pass as the picture. xAI's price list puts it at $0.08 per second of generated footage. The widely shared claim that it is "86 per cent cheaper than Sora" sets a 720p rate for Grok, as quoted by tech sites, against the price of Sora 2 Pro, OpenAI's most expensive video family. It therefore says nothing about what you would pay for comparable quality.

In short

  • xAI made the model available in its API on 3 June 2026, as a preview, with output up to 720p.
  • xAI's price list: $0.08 per second of generated video.
  • OpenAI's price list: $0.10 per second for sora-2 at 720p, and from $0.30 to $0.70 per second for sora-2-pro.
  • Sound (effects, ambience and speech) is rendered in the same pass as the picture.
  • In September 2026 the Artificial Analysis arena ranks the model sixth, with an Elo score of 1110.

What xAI actually released

According to xAI's own announcement, the model reached the API as a preview on 3 June 2026. It takes one starting frame plus a description of the motion and turns them into a smooth shot. xAI stresses two things: the detail and lighting of the input frame carry through into the video, and the look stays consistent across a series of shots, which matters when you cut several clips into one sequence. The same announcement caps resolution at 720p.

Model-gateway catalogues describe it a little more broadly. The model card in Vercel AI Gateway lists three output resolutions, 480p, 720p and 1080p, and states plainly that effects, ambient sound and speech are rendered in the same pass as the picture, in sync with the action. The gap between the maker's 720p ceiling and the intermediary's 1080p has not been explained anywhere, which is one more reason to treat this launch as a preview rather than a stable product.

Sound is generated in the same pass as the picture

In a typical AI video workflow, picture and sound are two separate steps. A model produces a silent shot, then a person or a second model adds music, effects and voice, and somebody lines them up with the frames. Every revised shot means repeating that editing work.

A model that emits picture and sound from a single pass removes the step. An impact, a footstep or a lip movement lands on the right frame because it was created with that frame, not stuck on afterwards. The practical difference is not that the sound is nicer. It is that one failed generation costs one generation, rather than a generation plus an hour of editing.

Sound generated in the same pass as the picture is not a gimmick. It is the difference between a clip and a finished scene.

What it really costs

Both vendors publish their price lists, so instead of repeating the marketing you can put them side by side. xAI's documentation gives Grok Imagine Video 1.5 a rate of $0.08 per second. OpenAI's API price list gives $0.10 per second for sora-2 at 720p, and $0.30, $0.50 and $0.70 per second for sora-2-pro at 720p, 1024p and 1080p respectively. These are list prices in US dollars as cited in the Polish original of this article, last revised 9 September 2026, so check the vendor pages before you set a budget.

Model and tierRate per second10-second clip
Grok Imagine Video 1.5 (xAI)$0.08$0.80
sora-2, 720p (OpenAI)$0.10$1.00
sora-2-pro, 720p$0.30$3.00
sora-2-pro, 1024p$0.50$5.00
sora-2-pro, 1080p$0.70$7.00

The clip column is our own multiplication of the list rate by ten seconds, not a figure supplied by either vendor. Keep that in mind, because in practice failed generations push the bill up: you pay for every one, including the ones you never use.

Where the 86 per cent figure came from

Tech sites covering the launch, gagadget among them, set the 720p rate they quoted for Grok, $4.20 per minute, against a Sora 2 Pro price of $30 per minute. For that pair the arithmetic is right: $4.20 is 14 per cent of $30, hence "86 per cent cheaper". The problem lies elsewhere. The comparison pits one vendor's 720p tier against the other vendor's most expensive model family, and gagadget does not say which resolution the $30-per-minute figure refers to.

At the same resolution the picture flips. The Artificial Analysis catalogue prices Grok's 720p tier at $8.40 per minute, or $0.14 per second, against $0.10 for sora-2 at the same resolution. The $4.20-per-minute rate behind the 86 per cent headline is exactly half the catalogue figure. Nobody has explained the discrepancy, so the only honest answer is to run the numbers on your own scenario, at the resolution you will actually use. It is the same discipline we described in our piece on tokenmaxxing and runaway AI budgets (in Polish).

What the blind-comparison arena says

An arena is a simple mechanism. A user is shown two clips generated from the same prompt, is not told which model made which, and picks the better one. Thousands of those choices add up to a ranking on the Elo scale familiar from chess. The upside is that brand reputation has no say in the verdict. The downside is that it measures preference, not correctness: a model can win with a striking frame while getting the physics of motion wrong. How rankings like this become a business in their own right is the subject of our piece on the arena that makes money comparing AI models (in Polish).

In June 2026 trade sites reported that Grok Imagine Video 1.5 had reached the top of an image-to-video ranking. gagadget does not say who runs that ranking, and separately mentions blind voting on DesignArena. In September 2026 the Artificial Analysis leaderboard, in its with-audio variant, has it in sixth place:

  1. Minimax H3 Max, Elo 1200
  2. Dreamina Seedance 2.0 (720p), Elo 1192
  3. MiniMax H3, Elo 1187
  4. Gemini Omni Flash, Elo 1181
  5. Wan 3.0, Elo 1176
  6. Grok Imagine Video 1.5, Elo 1110

The planning lesson matters more than the order. In video generation, first place has a shelf life measured in weeks, so do not build a process around one vendor just because it leads today.

What it means for a small business

The most realistic use is not a commercial but high-volume production of short variants. Eight versions of a six-second product shot add up to 48 seconds of footage, which by our calculation comes to about $3.84 at xAI's base rate (48 × $0.08). For that money you get material to test eight ideas for a Reel, rather than one shot that has to work first time. Variants only earn their keep once they are tested in front of an audience, organically or in paid placements. Alongside SEO, GEO and AEO, we run Google Ads campaigns, local campaigns in Google and campaigns in Meta Ads (see what we do).

The limits are just as concrete. 720p is enough for vertical social formats, but not for a big screen or for footage a customer will watch on a television. Before you publish anything, check the vendor's terms for what commercial use of generated footage is allowed, and label AI-generated content wherever the platform requires it. Model output is cheap raw material, not a finished campaign, and it is no substitute for content that keeps working after the ad budget stops.

The caveats

xAI itself calls this version a preview, so the model's behaviour, limits and pricing can change without notice. That is not a foundation for a production process with client deadlines attached.

We know of no methodology beyond the arena. xAI has not published a model card describing the training data, nor a repeatable quality measurement, so the whole case for the model's lead rests on user votes. Any internal comparisons from the vendor, should they appear, would be marketing material until someone independent repeats the measurement. It is the same reason AI assistants give little weight to what a company says about itself, as we explain in where AI citations come from.

We have not tested the model on client material. This article sets public price lists against a public ranking; it is not our own benchmark. Reports of a higher-resolution mode and of generation times measured in tens of seconds remain unconfirmed, in our view, until they appear in xAI's documentation.

How to test it on your own work in a week

  1. Pick one real use case. For example, a six-second animation of a product photo for Reels.
  2. Prepare a fair test set. Five starting frames and five motion descriptions, all written to the same template, so the comparison means something.
  3. Generate the same set twice. Once in Grok Imagine Video 1.5 and once in sora-2 at 720p, recording the cost and the time of every attempt.
  4. Review blind. Show both sets to three people outside the project, without the model names.
  5. Calculate the cost per usable clip, not per generation.

That last number, not a leaderboard position, decides whether the tool makes sense for your budget. Coinbase took a similar route when it cut its AI bills in half (in Polish), switching provider wherever the difference in quality was invisible to users.

The pace of xAI's releases is also tied to the capital position of Elon Musk's companies. We cover that background in our piece on the SpaceX stock-market debut (in Polish).

Common questions

How much does Grok Imagine Video 1.5 cost?

xAI's price list gives $0.08 per second of generated video. The Artificial Analysis catalogue prices the 720p tier at $8.40 per minute, which works out at $0.14 per second.

Is Grok Imagine Video 1.5 cheaper than Sora 2?

It depends on the tier. xAI's base rate is lower than sora-2's and well below sora-2-pro's, but at the same 720p resolution Grok's catalogue rate is sometimes quoted above sora-2's.

What does native audio mean?

That effects, ambient sound and speech are generated in the same pass as the picture and stay in sync with it, instead of being added in a separate editing step.

Is it the best video model on the market?

In June 2026 trade sites reported it in first place in an image-to-video ranking, but in September 2026 the Artificial Analysis leaderboard puts it sixth, with an Elo score of 1110.

Read next

Find out whether AI recommends your company.

Start with the free SEO and GEO audit, delivered in 5 working days. We check how the models describe your brand and hand back a prioritised list of changes.