Overview
Groq provides fast AI inference for openly-available models, with a focus on ultra-low-latency performance and compatibility with OpenAI endpoints.
Ultra-Low-Latency Inference
Groq's AI inference engine provides fast and reliable performance, with latency as low as 1ms.
OpenAI Endpoint Compatibility
Groq's API is compatible with OpenAI endpoints, making it easy to integrate with existing workflows and tools.
Support for Openly-Available Models
Groq supports a range of openly-available models, including Llama 3.1 and other models from leading AI research organizations.
Get started
- Open the official website and confirm the service is available in your region.
- Check the current plan, usage limits and terms for your intended use.
- Try a small task with sample data before committing to a paid plan.
Editorial note
Pricing checked: Sep 20, 2026. Free plan or free trial available with optional paid upgrades for higher limits.
Record updated: Sep 20, 2026
Suggest a correction