All models

gemini-3.5-flash

GoogleChatReasoningVision
Get your API key
gemini-3.5-flash

A balanced reasoning model for coding iterations and multi-step tasks

Gemini 3.5 Flash is Google's stable multimodal reasoning model, focused on coding agents, subtask collaboration, and long-running workflows that require continuous progress. It delivers analysis, code, and structured results in text, making it suitable for balancing reasoning capabilities and response efficiency through repeated planning, tool calls, and feedback checks, while also using image understanding for complex tasks.

GoogleModel brand
ChatModel type
Reasoning, visual understandingTask capabilities
STANDARD APIs · QUICK SETUP

Keep your SDK. Connect in minutes.

Point the Base URL to api.acedata.cloud, configure your platform API key and the model ID below, and use your compatible SDK or client.

API hostapi.acedata.cloud
modelgemini-3.5-flash
OpenAI Python SDK
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ACEDATACLOUD_API_KEY"],
    base_url="https://api.acedata.cloud/v1",
)
response = client.chat.completions.create(
    model="gemini-3.5-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Choose an available protocol for this model. OpenAI SDK uses a Base URL ending in /v1; Anthropic SDK uses the root URL. See each guide for protocol-specific parameters, tools and response formats.

Specifications and interface features

Clarify capacity, input and output, and invocation methods before choosing a model.

Native input limit
1,048,576 tokens
Native output limit
65,536 tokens
Native input and output
Text, image, video, audio, and PDF input; text output
Reasoning and result organization
Native support for Thinking, function calling, and structured output
Text and image invocation
Chat Completions: messages accept text and image_url content blocks
Generation limitations
Does not support image generation, audio generation, or the Live API

Native capacity and modality describe model capabilities; this platform's actual input formats, tool usage, and response methods vary by the selected invocation entry point.

Core Capabilities

Learn what gemini-3.5-flash can bring to your work.

Keep Coding Cycles Moving Forward

Gemini 3.5 Flash focuses not on writing all the code at once, but on adapting to a cycle of analyzing requirements, proposing changes, receiving test feedback, and continuing to refine. You can input code snippets, error logs, and constraints together to have it generate patch suggestions and testing ideas; tests are executed by tools provided by the application, making it easy to retain records of each review round.

Joint Analysis of Long Materials, Text, and Images

Million-scale native input capacity is suited to organizing lengthy code, documents, and task histories, while image understanding can supplement screenshots, charts, and interface information. Clearly specifying the items to compare, evidence to focus on, and delivery format in the prompt can keep the output centered on specific questions rather than merely compressing large amounts of material into a generic summary.

Integrate Reasoning into Business Processes

Tool workflows should be organized according to the tool definitions and result formats of the selected public interface. The model is responsible for planning, interpreting results, and generating call suggestions; querying, running code, and writing are completed by the execution environment provided by the application. Actual completion status should come from tool returns and verification records, not be judged solely from the model's description.

Applicable Scenarios

Start with specific tasks to find where the model can be effective.

Development Debugging and Refactoring

Input relevant functions, failing cases, logs, and expected behavior, and have the model first list possible causes, then propose modification plans and a regression testing checklist. After the application runs the tests, feed the results back into the next round, ultimately delivering reviewable modification suggestions, risk notes, and test items, suitable for development assistance requiring multiple iterations.

Document Review and Difference Extraction

Suitable for reviewing specifications or technical reports in stages: first extract constraints, then compare version differences, and finally create an action list. Require each suggestion to correspond to a location in the material and identify conflicts that remain unexplained; this type of continuous work better demonstrates the ongoing-task positioning of 3.5 Flash, rather than merely generating a one-time generalized summary.

Screenshot-Driven Task Decomposition

Combine product screenshots, charts, and textual requirements into text-and-image messages, allowing the model to identify key content and break down to-dos, such as organizing interface issues, explaining chart changes, or generating acceptance steps. Output is organized as text or JSON for easy entry into a ticketing system; screenshot analysis and actual interface operation are different tasks, and clicks will not be completed automatically.

How to choose this model

Choose based on task complexity, input materials, and expected results.

Choosing between the stable and preview versions

gemini-3.5-flash is the stable version invocation name, while gemini-3-flash-preview is the preview version name; clearly specify the target when integrating. Existing applications that have established workflows around the output format, tool loops, and test sets of 3.5 Flash can continue evaluating with this model. When comparing newer Flash models, validate them on the same tasks rather than judging migration benefits solely by version number.

Choose the model tier by task difficulty

When multimodal analysis, coding iteration, and continuous planning are needed, 3.5 Flash is a balanced choice. For simple classification, fixed-field extraction, and retryable batch processing, compare Flash Lite; for difficult reasoning and complex analysis, compare Pro. Focus on task success rate, number of tool rounds, and total output volume, and do not treat the length of a single response as a substitute metric for reasoning quality.

Start with a specific task

Based on the characteristics of gemini-3.5-flash, first validate small tasks whose results can be checked.

01

Deliver multi-step tasks in stages

You can ask directly: Organize materials in stages according to business goals, compare options, and form a conclusion. For each stage, list the evidence used and information still needed, and output the inputs required for the next stage.

02

Prepare inputs that support decisions

Define stage goals and tool responsibilities; use intermediate results and final deliverables to check whether ongoing tasks have deviated from requirements.

03

Then integrate it into your workflow

Use the full model ID gemini-3.5-flash, first confirm the public request format and available parameters on the API page, then connect your application. Retain result parsing, exception handling, and relevant evidence, and use the same set of real samples to evaluate whether it is suitable for continued use.

Usage boundaries

Before formal use, understand output quality and capability scope.

  • This is a text-output model. It can understand multimodal materials, but it does not generate images, speech, or real-time bidirectional audio and video. Consider audio analysis and audio generation separately; when you need finished images, voice-overs, or real-time conversations, choose the corresponding generation or real-time model.
  • Long inputs do not mean that all content will be used with equal accuracy. Code repositories and long documents should include objectives, key locations, and acceptance criteria; when generating longer answers, also reserve a reasoning budget and avoid setting output limits too low, which may result in an empty or incomplete final response.
  • Native code execution and preview Computer use capabilities do not mean that ordinary chat requests will automatically run programs or operate a computer.

Frequently Asked Questions

Answers to common questions when using gemini-3.5-flash.

How do I choose between the two Gemini 3.5 Flash endpoints?

Use Chat Completions, set model to gemini-3.5-flash, organize text and images in messages, and read responses from choices; for streaming interactions, handle incremental updates as described in the documentation. The capabilities and message format of native Generate Content should be verified independently; do not infer that the platform exposes all features from the vendor's native specification.

How can I make Gemini 3.5 Flash analyze images?

In Chat Completions, write message content as an array of content blocks, combining text and image_url, and place an accessible image URL or Base64 data URI in image_url.url. The text should specify the objects of interest and output requirements, such as extracting fields, comparing differences, or explaining charts.

Can I have it read PDFs?

Prepare the document text, table data, or clear page screenshots relevant to the question, and specify whether you need a summary, comparison, or particular information extracted. Submit content in formats supported by the selected public API; a PDF URL cannot be used as image_url. Request that results retain original-text locations, field evidence, and unconfirmed items, and verify key numbers against the source material.

Why are tokens consumed but no final text is returned?

Gemini 3.5 Flash uses reasoning tokens, and an output budget that is too small may be exhausted before the final answer is formed. The usage guide recommends setting max_tokens to 512 or higher; complex tasks should allow even more space, and you should use finish_reason, usage, and the actual text to determine whether the budget needs to be increased.

Is it the same invocation name as gemini-3-flash-preview?

No. Google lists gemini-3.5-flash as Stable and gemini-3-flash-preview as Preview. To call this model, use the full ID gemini-3.5-flash; when replacing models, especially retest structured outputs, tool loops, and long-task performance rather than putting it into use immediately after only changing the name.