Prompting and workflow

Image to 3D model

Short answer

A reference image tells an AI CAD engine what shape you mean. It cannot tell it what size that shape is. A photograph contains no scale, so dimensions still have to come from your prompt. Use the image for form and the text for numbers, and you get a model that is both the right shape and the right size.

A photograph contains no scale. That single fact explains almost everything about what image-to-3D can and cannot do, and it is the thing most tools in this space are quietest about.

Point a camera at a bracket and the image records shape, proportion and arrangement. It does not record that the part is 60 mm long, because the same pixels would result from a part 6 mm long photographed closer, or a part 600 mm long photographed further away. No amount of model capability recovers a number that was never captured.

This is not a limitation to work around. It is a division of labour. The image supplies the form; you supply the numbers; together they produce a model that is both the right shape and the right size.

What a reference image actually contributes

Attach an image in TextoCAD and it goes to the model alongside your prompt, resolving the ambiguity that words carry badly. Some things are genuinely tedious to describe and instant to show:

  • Overall formwhen the part is not a name you can reach for. "A sort of tapered bracket with a stepped foot" is three attempts at a sentence and one glance at a picture.
  • Arrangement. Which side the flange is on, how the cutouts sit relative to the mounting holes, which way the hook faces.
  • Proportion. Not absolute size, but the relationship between parts: that the base is roughly twice the height, that the slot runs most of the length.
  • Style on decorative work, where a described shape and an intended shape can diverge a long way.

What it does not contribute is any dimension whatsoever. That stays your job, and it is why the prompt still matters.

Taking a usable reference image

The difference between a good and a poor reference is mostly about removing things the model has to interpret.

  • One part, filling the frame. A workbench with six objects on it forces a guess about which one you mean.
  • Plain background. Sheet of paper, plain worktop, anything without pattern. Contrast between part and background does most of the work.
  • Square on, not three-quarter. A straight-on view of the face that matters beats a more attractive angled shot. If two faces matter, take two photographs.
  • Flat, even light. Hard shadows read as features. Diffuse daylight is ideal, and a direct flash is the worst option.
  • Put a ruler in the frame. It will not be measured for you, but it keeps you honest when you check the result against the original.

Sketches often work better than photographs

A hand sketch has an advantage a photograph cannot match: it contains only what matters. No background, no lighting, no reflections, no perspective distortion. Every line you drew is a line you meant.

The strongest version is an engineering-style sketch with the dimensions written on it. Draw the part roughly to proportion, mark the overall sizes, hole diameters and thicknesses directly on the drawing, photograph it flat, and attach it. You have now communicated shape and numbers in a single image, which is exactly the split this whole page is about.

This works well for parts you are copying from something physical. Measure the original with calipers, sketch it, annotate it, generate it.

How this differs from photogrammetry and mesh reconstruction

Two other technologies get grouped under "image to 3D" and they do a different job.

Photogrammetry reconstructs a surface from many overlapping photographs taken around an object. With enough coverage it produces an accurate mesh of what is really there, and it is the right tool for capturing a statue, a rock, a face. It needs dozens of images, it struggles with shiny and featureless surfaces, and it gives you a scan.

Single-image mesh generation takes one photo and infers a plausible three-dimensional shape, guessing the parts it cannot see. Impressive for concept work and game assets. Not a measurement.

Both produce meshes: triangles with no editable dimensions and no STEP export. If your goal is a decorative object or a digital asset, they are the right tools. If the part has to bolt to something, the thing you need is a solid, and AI CAD generators compared covers that distinction in full.

The practical workflow

Image plus stated dimensions, then correct on the sliders. Concretely:

Photographed bracket

Match the shape of the bracket in the attached image. It is an L bracket with a 70 mm vertical leg and a 45 mm horizontal leg, 25 mm wide and 4 mm thick, with two 4.5 mm holes in each leg spaced 25 mm apart, and a 4 mm fillet at the inside corner.

The image fixes the shape and the orientation. Every number still comes from the sentence.

Annotated sketch

Build the part shown in the attached sketch, using the dimensions marked on it. Overall 120 by 40 mm, 6 mm thick, with the slot 60 mm long and 8 mm wide running along the centreline, and 3 mm fillets on the outer corners.

Restating the marked dimensions in the prompt removes any dependence on the handwriting being legible.

Replacement for a measured original

Match the profile of the knob in the attached photo. It is 32 mm in diameter and 18 mm tall with twelve grip flutes around the circumference, a 6.2 mm D-shaped shaft socket 12 mm deep, and a 1 mm chamfer on the top edge.

Measure the original with calipers first. The photo cannot supply the socket size and getting it wrong is the one thing that makes the part useless.

Why results disappoint, and how to fix it

Nearly every poor outcome traces to the same cause: the image was treated as a replacement for the prompt rather than an addition to it. An image with a one-word instruction gives the engine a shape and no requirements, so it invents the requirements. They will not be yours.

The rest, in rough order of frequency:

  • No dimensions given. The most common by a wide margin. Say the size.
  • Busy background. Reshoot against plain paper. It takes twenty seconds and changes the result.
  • Angled photograph. Perspective makes a rectangle a trapezium. Shoot square on.
  • Hidden features. A hole on the far side is not in the image. Describe it or photograph that side too.

When the shape is right and the numbers are not, do not re-prompt. Move the sliders. Correcting a thickness on a slider is faster than describing it again, and it does not spend a generation.

Where to go next

The prompt half of this is covered properly in CAD prompt examples, and how to convert text to CAD walks the whole workflow. If the part is destined for a printer, design for 3D printing has the clearances to state in the prompt. The gallery shows the prompt behind every model, which is the fastest way to calibrate how much detail is enough.

Step by step

  1. 1

    Choose or take a usable image

    One part, filling the frame, against a plain background, photographed square on rather than at an angle. A straight-on view of the face you care about is worth more than an artistic three-quarter shot.

  2. 2

    Measure the part yourself

    Take calipers to the real object, or a ruler to the drawing. The image cannot supply these numbers and the engine will not invent them correctly. Overall size, thickness, and any hole diameter are the minimum.

  3. 3

    Write the prompt with the numbers in it

    Attach the image and describe the part as you normally would, stating every dimension you measured. The image resolves ambiguity about shape; the text fixes the scale.

  4. 4

    Correct on the sliders

    Compare the generated dimensions against your measurements and adjust. This is faster than re-prompting and does not consume a generation.

Frequently asked questions

Can AI turn a photo into a CAD model?

It can turn a photo into an interpretation of the shape. It cannot recover true dimensions from a single image, because a photograph has no scale in it. Tools that claim one-click photo to CAD are either producing a mesh or quietly guessing the size.

What makes a good reference image?

One part, plain background, shot straight on rather than at an angle, and well lit with no strong shadows across the features. A photograph of the part next to a ruler is more useful than a better-composed photo without one.

Can I use a hand sketch instead of a photo?

Yes, and often it works better. A sketch shows only what matters, with no background or lighting to interpret. Write the dimensions directly onto the sketch before photographing it and you have communicated both the shape and the numbers in one image.

How is this different from photogrammetry?

Photogrammetry reconstructs a surface from many overlapping photographs and produces a mesh, which is the right tool for capturing an existing organic object. It gives you a scan, not a designed part: no editable dimensions and no STEP file.

Does the image replace the prompt?

No, and treating it as a replacement is the usual reason results disappoint. The image is context, the prompt is the specification. An image with a one-word prompt gives the engine a shape and no requirements.

Try an image with your measurements

Attach a straight-on photo, state the dimensions you measured, and adjust the result on the sliders.