How to create an AI prompt from a photo?

Updated 2026-08-02AI-assisted draft · citations disclosedPart of the 1,478-question editorial index· prompt engineering and LLM · Source & maintenance record
Short answer

Upload the photo to a multimodal model and ask for a structured image-generation prompt, separating what is visibly observed from what is only inferred. Request subject, composition, lighting, color, texture, mood, camera-like cues and constraints, then adapt the draft to your target generator and iterate one change at a time. A text prompt will not recover hidden metadata or guarantee the same person, place or image; get consent and check the tool's privacy and usage terms before uploading or publishing.

Why — the first-principles explanation

Turning a photo into a prompt is visual analysis followed by creative specification, not a perfect reverse button. A vision model converts pixels into a description; an image generator converts a description and optional reference images into a new image. Text is a lossy representation, so the generated image can share composition, lighting or mood without preserving the original face, exact location, camera metadata or unrecorded intent. If you need the subject or layout to remain close, use an image-to-image or reference-image workflow instead of text alone.

The most useful prompt is structured around controllable decisions. Ask the model to identify the main subject and action, framing and camera angle, foreground/midground/background, light direction and quality, palette and contrast, texture, mood, environment, aspect ratio and the purpose of the new image. Ask it to label each detail as observed, likely, or unknown. A guessed “85mm lens” is not a fact from the pixels; it is a hypothesis that may be useful as a visual cue.

Target matters. A conversational image editor may understand natural-language revisions, while a developer API may expose image inputs, detail settings, aspect ratios, reference images or structured outputs. Midjourney, Stable Diffusion, Imagen, Gemini or another generator can interpret the same words differently. Start with a portable description, then add only the target tool's documented syntax or parameters. Do not carry a random negative-prompt list or obsolete model flag from one tool into another.

Iteration beats a giant prompt. Generate a baseline, compare it with the reference, change one high-impact variable—framing, subject, light, palette or level of abstraction—and keep the prompt and seed or generation settings when the tool supports them. For complex scenes, break the task into steps or use a reference image; for text-heavy assets, inspect spelling and layout manually. The model may miss small text, rotated content, reflections, counts or fine spatial relationships, so use a high-detail input or a crop when your tool supports it.

Rights and privacy are part of the workflow. Do not upload an identifiable person's photo without an appropriate basis or consent, and do not assume that a prompt naming a living artist or a copyrighted character gives you permission to reproduce it commercially. Google's current Gemini image guidance tells users to respect others' copyright and privacy rights and notes that availability and age requirements vary. Keep the original, the generated output, the prompt and the tool/version record so you can explain what was transformed and under which terms.

An example that makes it click

Suppose you have your own photo of a ceramic cup on a kitchen table and want a product-style image. Ask the vision model for a JSON-like brief: subject and materials; camera angle and crop; window direction and softness; background objects; palette and surface texture; observed text or logos; unknowns; and three changes for a square product shot. Use only the observed logo or remove it intentionally, choose a target image tool, generate a baseline, then change the background from a busy kitchen to a neutral warm-gray sweep. The prompt gives you a repeatable starting point without pretending it recovered the original camera settings.

How to do it

  1. Confirm that you have the right to upload and transform the photo. Obtain appropriate consent for identifiable people, remove unnecessary private details and check the destination tool's retention, training and sharing terms.
  2. Define the goal: describe the photo, make a similar new image, edit the same image, preserve a person or object, change the background, or create a commercial asset. Different goals need different workflows.
  3. Choose a multimodal model that accepts image input and a target generator or editor. Record the product, account, model, date and region because capabilities and limits change.
  4. Upload a clear, correctly oriented image. If fine text, faces, small objects or spatial placement matter, crop the relevant area or select the tool's higher-detail input option when available.
  5. Ask for a structured analysis, not just a caption. Request subject/action, composition, foreground/midground/background, lighting, palette, texture, mood, style cues, aspect ratio and constraints.
  6. Require uncertainty labels. Have the model separate directly visible observations from inferred camera, lens, time, identity, location or artistic-intent guesses; delete guesses that do not help the target task.
  7. Ask for two outputs: a portable natural-language brief and a target-tool prompt or parameter block. Keep brand names, real people, copyrighted characters and living-artist style references only when you have a legitimate reason and the tool permits them.
  8. Generate a baseline, compare it with the reference and change one variable at a time. Save the prompt, reference images, settings, output and reason for each revision so the result is reproducible.
  9. Switch to image-to-image, inpainting or a reference-image feature when text cannot preserve identity, pose, geometry, typography or a particular object. Inspect spelling, faces, hands, counts and logos manually.
  10. Before sharing or selling, verify consent, licenses, disclosure or watermark requirements, privacy, safety and the generator's current terms. Keep the source and transformation record instead of presenting a synthetic image as an untouched photograph.

Key facts

Infographic: How to create an AI prompt from a photo — short answer and key facts
Visual summary — How to create an AI prompt from a photo?

Turn a reference photo into a repeatable, rights-aware workflow

Separate observation from inference, choose text or reference-image controls for the real goal, then iterate with documented prompts, consent and the target generator's current terms.

▶ The 60-second explainer (script)

How do you create an AI prompt from a photo? Treat it as two steps: analyze the image, then specify a new image. Upload the photo to a multimodal model and ask for a structured brief covering subject and action, framing, foreground and background, light, palette, texture, mood, aspect ratio and constraints. Ask it to label what is observed versus inferred—an assumed lens or location is not a fact. Adapt the brief to your target generator, create a baseline and change one variable at a time. If you need the same person, pose, geometry or lettering, use a reference-image, image-to-image or inpainting feature instead of text alone. Before uploading, check consent, privacy and the service terms. Google’s guidance recommends specific context, clear constraints, iteration and human review; its image help also warns that generated results can be inaccurate and that availability varies by country and age. Keep the prompt and source record so the transformation is repeatable and honest.

What authoritative sources say

Google AI for Developers — Prompting strategiesofficial — Google's prompting strategies recommend few-shot examples, consistent structure, relevant context, breaking complex prompts into steps and iterating; its multimodal guidance says text, images, audio and video should be referenced clearly in instructions. source ↗
Google AI for Developers — Visionofficial — Google's vision documentation explains image inputs, supported formats, detail and resolution trade-offs, and cautions that clear, correctly oriented images help visual understanding while small text and fine details can be difficult. source ↗
Google AI for Developers — Image generationofficial — Google's image-generation guidance recommends specific visual detail, context and intent, iterative refinement, step-by-step instructions and positive descriptions of desired outcomes; it also documents reference-image and image-editing workflows. source ↗
Google Gemini Help — Generate and edit imagesofficial — Google Gemini Apps supports uploading a photo to ask for edits or combine multiple reference images, and its help page says users should respect copyright and privacy rights and check current country, language and age availability. source ↗
OpenAI API — Images and visionofficial — OpenAI's image documentation describes using images as model inputs, image detail levels and limitations such as small text, rotation, counting and resizing; the exact behavior is model-specific. source ↗

People also ask

Can an AI recreate the exact photo from a prompt?

Usually not. Text loses pixel-level identity, geometry, metadata and context. Use the original image as a reference or choose image-to-image, inpainting or an editing workflow when preserving the subject or layout matters.

Do I need a special photo-to-prompt tool?

No. A multimodal chat or image tool that accepts your photo can produce a structured draft. A dedicated wrapper may add convenience, but compare its upload, retention, export and billing terms before sending private images.

What should I ask the model to include?

Ask for subject and action, composition and framing, foreground/midground/background, light direction and quality, color and contrast, texture, mood, setting, aspect ratio, target use and constraints. Ask it to mark inferred details separately.

Can the model tell me the exact camera and lens?

Not reliably from pixels alone. It can suggest a lens or lighting description that may produce a similar look, but label it as an inference unless you have the original metadata or another source.

How do I preserve a person's face or an object's shape?

Use a reference-image, image-to-image, edit or inpainting feature with the controls your generator documents. A text description can guide identity or shape but cannot guarantee it; obtain appropriate consent for identifiable people.

Should I use negative prompts?

Only when the target generator supports and benefits from them. First state the desired scene positively and add a short, testable exclusion list; parameter syntax and how negative prompts are weighted vary by tool.

Why does my generated image not match the reference?

The prompt may omit a high-impact detail, the model may infer incorrectly, the reference may be too small or blurry, or the target tool may interpret the wording differently. Crop important areas, label uncertainty, change one variable and use a reference workflow when needed.

Can I turn someone else's photo into a prompt?

Only when you have an appropriate right or permission to upload and transform it. Avoid exposing private faces or details, review the service terms and do not imply that a synthetic result is an untouched photograph.

Can I ask for a living artist's style?

The legal and platform treatment of style imitation varies. A safer brief describes observable visual properties—medium, palette, composition, line quality and lighting—without using a living artist's name, and you should check the target tool's current policy.

Which model creates the best prompt from a photo?

There is no permanent winner. Compare the same reference and structured rubric across tools, including observation accuracy, uncertainty labels, target-generator fit, privacy terms, cost and how well the final image matches your goal.

Can I sell an image generated from a photo?

That depends on the source photo, identifiable people, third-party marks, the generator's terms and the law where you use it. Keep proof of permission and provenance, read the current commercial-use terms and get professional advice for a material commercial project.

The same question, asked other ways

This page answers one intent expressed in 2 phrasings. How the index is organized →

Related questions