Best text to image generator

posted 9 months ago

I have used several different generators. What they all seem to have in common is that they don’t always display what I am asking for. Example: if I am looking for a person in jeans and t-shirt, I will get images of a person wear things totally different clothing and it isn’t consistent. Another example is if I want a full body picture, that command seems to be ignored giving just waist up or just below the waist. Same goes if I ask for side views or back views. Sometimes they work. Sometimes they don’t. More often they don’t. I have also seen that none of the negative requests seem to actually work. If I ask for pictures of people and don’t want them using cell phones or no tattoos, like magic they have cell phones. Some have tattoos. I have noticed this in every single generator I have used. Am I asking for things the wrong way or is the AI doing whatever it wants and not paying attention to my actual request?

Thanks

Sort:

Hot Top Controversial New Old

[ - ]

vibinya@lemmy.world

16 points

9 months ago

My favorite has been locally hosting Automatic1111’s UI. The setup process was super easy and you can get great checkpoints and models on Civitai. This gives me complete control over the models and the generation process. I think it’s an expectation thing as well. Learning how to write the correct prompt, adjust the right settings for the loaded checkpoint, and running enough iterations to get what you’re looking for can take a bit of patience and time. It may be worth learning how the AI actually ‘draws’ things to adjust how you’re interacting with it and writing prompts. There’s actually A LOT of control you gain by locally hosting - controlNet, LORA, checkpoint merging, etc. Definitely look up guides on prompt writing and learn about weights, order, and how negative prompts actually influence generation.

permalink

report

[ - ]

EdgeRunner@lemmy.dbzer0.com

2 points

9 months ago

Ive started with stablediffusion_webui, i feel you !!

permalink

report

parent

[ - ]

rickdg@lemmy.world

12 points

9 months ago

Can you give an example of a complete prompt? Are you using Dall-E, Midjourney, Stable Diffusion…?

It seems that all models need to have prompts crafted specifically for them and you need to follow-up with corrections. The follow-up is critical for pretty much anything these LMMs output.

permalink

report

[ - ]

Ragdoll X@lemmy.world

4 points

9 months ago

Image-to-image also helps a lot with SD. Even some roughly-drawn blobs can be the difference between the image almost matching what you had in mind vs. looking exactly how you intended.

permalink

report

parent

[ - ]

BlueÆther@no.lastname.nz

1 point

9 months ago

I just cant get img2img on SD to work for me to get images that are what I want(A1111 front end)

permalink

report

parent

[ - ]

silas@programming.dev

6 points

9 months ago

Talking to a text-to-image model is kinda like meeting someone from a different generation and culture that only half knows your language. You have to spend time with them to be able to communicate with them better and understand the “generational and cultural differences” so to speak.

Try checking out PromptHero or Civit.ai to see what prompts people are using to generate certain things.

Also, most text-to-image models are not made to be conversational and will work better if your prompts are similar to what you’d type in when searching for a photo on Google Images. For example, instead of a command like “Generate a photo for me of a…”, do “Disposable camera portrait photo, from the side, backlight…”

permalink

report

[ - ]

EdgeRunner@lemmy.dbzer0.com

5 points

9 months ago

Its time to promote, https://lemmy.dbzer0.com/c/stable_diffusion_art.

Very helpfull and relaxing,

permalink

report

[ - ]

CommunityLinkFixerBot@lemmings.worldB

9 points

9 months ago

Hi there! Looks like you linked to a Lemmy community using a URL instead of its name, which doesn’t work well for people on different instances. Try fixing it like this: !stable_diffusion_art@lemmy.dbzer0.com

permalink

report

parent

[ - ]

simple@lemmy.world

3 points

9 months ago

Dall-E 3 is the easiest to use and usually understand prompts the best. You can use it for free via Bing Image Editor.

permalink

report

Technology

!technology@lemmy.world

Create post

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

Community stats

18K
Monthly active users
11K
Posts
505K
Comments

Our Rules

Approved Bots

Community stats

Community moderators