Gemini 1.5 Pro and Gemini 1.5 Flash prices dropped down, which is 50% cheaper compared with ChatGP- 4o, but with more updated and powerful models! Check the latest price and updates below:
Starting October 1, 2024, Gemini 1.5 Pro will see significant price reductions: a 64% reduction in input tokens, 52% in output tokens, and 64% in incremental cached tokens. As part of the strongest 1.5 series model, these reductions, combined with context caching, continue to drive down the cost of building with Gemini, making it more accessible and cost-effective for users.
The New Gemini 1.5 Flash-8B is ready with a lower price
Gemini 1.5 Flash-8B, the latest Gemini Flash variant, is production-ready and comes with:
- 50% lower price (compared to 1.5 Flash)
- 2x higher rate limits (compared to 1.5 Flash)
- Lower latency on small prompts (compared to 1.5 Flash)
With the stable release of Gemini 1.5 Flash-8B, Google is announcing the lowest cost per intelligence of any Gemini model ever since launching:
- $0.0375 per 1 million input tokens on prompts <128K
- $0.15 per 1 million output tokens on prompts <128K
- $0.01 per 1 million tokens on cached prompts <128K
*The billing is effective from 14 Oct 2024.
Gemini 1.5 and ChatGPT-4o Pricing Comparison
| Gemini 1.5 Pro | GPT-4o | Gemini 1.5 Flash-8B | GPT-4o mini | |
| Input Cost of input data provided to the model. | US$2.50 Per million tokens | US$5.00 Per million tokens | US$0.075 Per million tokens | US$0.15 Per million tokens |
| Output Cost of output tokens generated by the model. | US$10.00 Per million tokens | US$15.00 Per million tokens | US$0.30 Per million tokens | US$0.60 Per million tokens |
Cheaper But Most Powerful Ever
Google is releasing two updated production-ready Gemini models: Gemini-1.5-Pro-002 and Gemini-1.5-Flash-002. These latest models build on the recent experimental release, offering significant improvements to the Gemini 1.5 models announced at Google I/O in May. Developers can now access these advanced models for free through Google AI Studio and the Gemini API. For larger enterprises and Google Cloud customers, the models are also available on Vertex AI, ensuring accessibility for a wide range of users.
The Gemini 1.5 series models offer enhanced performance across various tasks, including text, code, and multimodal applications. These models excel in handling complex challenges such as synthesizing information from lengthy PDFs, analyzing large codebases, and generating content from long videos. With recent updates, the 1.5 Pro and Flash models are now more efficient, delivering faster and more cost-effective results. Notable improvements include a 7% increase in MMLU-Pro performance, a 20% boost in math benchmarks, and 2-7% gains in vision and code generation tasks. These enhancements make the models more robust for diverse use cases.
To make building with Gemini more accessible, Google is increasing rate limits for the paid tiers.
- For 1.5 Flash, the rate limit is now 2,000 RPM.
- For 1.5 Pro, it is raised to 1,000 RPM, up from 1,000 and 360 RPM
A further increase is expected to Gemini API rate limits, and Google has reduced latency and increased output speed, with 1.5 Flash now delivering 2x faster outputs and 3x less latency, opening doors to new use cases.
Conclusion
The recent updates to Gemini 1.5 Pro and Flash models, alongside substantial price reductions, provide developers with more powerful, cost-effective AI tools. These enhancements, including faster outputs, lower latency, and increased rate limits, make it easier to scale your AI projects. With improved performance across various tasks like math, code, and vision, Gemini is ready to unlock new possibilities for your business.
Ready to experience the power of Gemini? Contact us today for a demo to see how these updates can benefit your development needs!






