For gemini-3.1-flash-lite on Vertex AI v1 in the US multi-region, I need an authoritative billing bound for one text-only generateContent request.
Settings: candidateCount=1, maxOutputTokens=8192, thinkingConfig.thinkingLevel=MINIMAL, thinkingConfig.includeThoughts=false, structured JSON output, and no tools, grounding, cache or multimodal input.
Does maxOutputTokens guarantee candidatesTokenCount + thoughtsTokenCount <= 8192 for this exact model and endpoint? If not, what is the separately enforced maximum billable thinking-token count? Does MINIMAL establish a numerical maximum, or only a behavioral setting?
Please also distinguish billing for client cancellation from API error responses. The general model, thinking and pricing pages do not clearly establish this exact request guarantee. Please cite the applicable specification or have the product team confirm the bound.
This is a pre-use cost-control question, not a request to change any account or service.