I’m building an automated workflow to intelligently crop decorative wallpaper images based on customer wall dimensions (width × height in cm). The catalog has 7000+ images with very different subjects (flowers, leaves, geometric patterns, figures, etc.).
My current approach:
AI Agent with OpenAI Vision analyzes the image and returns crop coordinates as percentages (left, right, focal_y)
A JavaScript node converts percentages to pixels based on image dimensions and wall size
A Python/Pillow node executes the actual crop
The problem is that GPT Vision estimates coordinates visually and is sometimes imprecise — subjects like leaves or flowers end up slightly cut at the edges.
What I already tried:
Setting temperature: 0 for determinism
Detailed prompt with design principles and examples
focal_y for vertical positioning
My question: What would be the best approach to improve precision on a large and varied catalog?
Refine the prompt further?
Add OpenCV in the Python node to detect precise object boundaries after GPT gives the approximate area?
Also, prompts are everything in AI agents, so you have to check your prompts here:
Your project idea is awesome, but 7k images are very concerning. Instead, why don’t you let AI generate the image based on the user requirement? I think that would be cheaper than actually picking an image from that huge catalogue.
Me topé con tu hilo mientras investigaba el mismo problema desde un ángulo diferente — usar Claude (no GPT) para obtener coordenadas de recorte para fotografía deportiva editorial, y luego aplicar el recorte a través de ImageMagick en n8n. Misma arquitectura, misma frustración: el modelo estima visualmente y ocasionalmente pierde el sujeto.
Unas cuantas preguntas si has tenido tiempo de iterar desde abril:
¿Terminaste agregando OpenCV para refinamiento de bordes después del paso del LLM, o la ingeniería de prompts te llevó lo suficientemente lejos?
¿Probaste cambiar a un modelo diferente (Gemini, Claude) o a una API dedicada de recorte como Imagga?
¿Algún aprendizaje sobre cómo estructuraste la salida de coordenadas — porcentajes vs. píxeles, caja delimitadora vs. punto focal?
Estoy feliz de compartir mis iteraciones de prompts si eso te es útil a cambio.