llama-3.2-90b-vision-instruct by deepinfra - AI Model Details, Pricing, and Performance Metrics

meta
llama-3.2-90b-vision-instruct
meta

llama-3.2-90b-vision-instruct

completions
bydeepinfra

The Llama 90B Vision model is a top-tier, 90-billion-parameter multimodal model designed for the most challenging visual reasoning and language tasks. It offers unparalleled accuracy in image captioning, visual question answering, and advanced image-text comprehension. Pre-trained on vast multimodal datasets and fine-tuned with human feedback, the Llama 90B Vision is engineered to handle the most demanding image-based AI tasks. This model is perfect for industries requiring cutting-edge multimodal AI capabilities, particularly those dealing with complex, real-time visual and textual analysis. Click here for the [original model card](https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD_VISION.md). Usage of this model is subject to [Meta's Acceptable Use Policy](https://www.llama.com/llama3/use-policy/).

Context
32768
Input
$0.35 / 1M tokens
Output
$0.4 / 1M tokens
Accepts: text, image
Returns: text

Access llama-3.2-90b-vision-instruct through LangDB AI Gateway

Recommended

Integrate with meta's llama-3.2-90b-vision-instruct and 250+ other models through a unified API. Monitor usage, control costs, and enhance security.

Unified API
Cost Optimization
Enterprise Security
Get Started Now

Free tier available • No credit card required

Instant Setup
99.9% Uptime
10,000+Monthly Requests

Code Examples

Integration samples and API usage