ImagineArt 2.0 Edit · Model case study 02

Building ImagineArt 2.0 Edit

The editing model was designed around direction instead of a layer stack: bring up to four reference images, describe the intended frame, and preserve what matters while changing what does not.

I led research with our ML team for ImagineArt 2.0 Edit, the image-grounded counterpart to the 2.0 text-to-image model. The project asked a harder question than “can the model edit an image?” We wanted one system to reason across people, products, garments, scenes, poses and styles while preserving the identity and structure that made the input valuable.

From instruction to composed frame

Traditional editing workflows make the user manage masks, layers, transforms and lighting corrections. A generative editing model can compress those operations into a natural-language direction, but only if it understands the role of every reference image.

ImagineArt 2.0 Edit accepts as many as four references in one request. A user can provide a subject, another person, a garment and an environment, then direct how they should appear together. The model must preserve local evidence—face, fabric, product geometry—while creating a globally coherent frame.

My research scope

My role covered the model programme from capability definition through production readiness. I led the team working on multi-reference conditioning, identity preservation, aspect-ratio behavior and the evaluation needed to compare editing quality across very different use cases.

The model had to support a connected set of workflows:

  • multi-subject and scene composition;
  • identity-preserving changes to setting, lighting and wardrobe;
  • outfit swap and product placement;
  • style transfer without losing structural geometry;
  • background replacement and subject insertion;
  • pose and expression transfer.

Preserving identity under change

Identity preservation is not a single similarity score. A recognisable face can still feel wrong when body proportion, age cues, hair or expression drift. Product identity has its own requirements: shape, material, hardware and branding need to survive contact with a newly generated scene.

We treated preservation and edit strength as a controlled tension. If conditioning is too rigid, the new frame looks collaged. If it is too weak, the subject becomes generic. Evaluation therefore paired reference fidelity with lighting integration, pose plausibility and scene coherence.

An editing model succeeds when the result feels newly photographed, while the people and products still feel unquestionably the same.

Multi-image composition

Multiple references create an assignment problem. The instruction must establish who goes where, which garment belongs to which person and whether an image represents content, style or setting. ImagineArt 2.0 Edit was developed so these roles can be expressed in ordinary direction rather than a specialist syntax.

Composition also requires a common visual world. Shadows, camera angle, depth and color response should agree across inputs. The released model is designed to integrate those decisions in one generation instead of relying on separate compositing and relighting passes.

Native canvases and production output

Editing models often force inputs through a square canvas and crop them back afterwards. That breaks composition and wastes the information in wide or vertical references. We developed native aspect-ratio behavior so the first image can define the frame or the user can choose a target ratio explicitly.

The production release supports eight named aspect ratios plus automatic matching, with output up to 2048 pixels. These behaviors are documented on the official ImagineArt 2.0 Edit product page and exposed through the ImagineArt API.

Evaluating edits as instructions

Image quality alone cannot measure whether an edit worked. We built evaluation around the contract expressed by the instruction: what must stay, what must move and what must change. Each capability needs targeted cases, followed by mixed cases that expose conflicts between preservation and transformation.

For example, outfit replacement tests garment fidelity, body preservation, contact, occlusion and lighting at once. Subject insertion tests identity, scale, perspective, shadow and environmental integration. A useful review records the failure dimension rather than collapsing every judgment into one preference score.

One model, a wider creative surface

The final model turns a broad family of image-editing tasks into one directed workflow. Creators can cast subjects, change wardrobes, place products, transfer style and re-stage scenes without moving between separate tools.

For the research team, the larger achievement was unification: one image-grounded model, one production interface and a shared evaluation language for control. That is the foundation for creative systems that respond less like a filter and more like a collaborator.

Muhammad Ahmed Ghani

About Muhammad Ahmed Ghani

Muhammad Ahmed Ghani is AI/ML Lead and Lead Researcher at ImagineArt. He currently works in Islamabad, Pakistan, and is originally from Lahore. View his profile and model portfolio.