

https://civitai.com/models/214823?modelVersionId=242025
Welcome to try out the open-source GPT4V-Image-Captioner, developed by my friend and me. It offers a one-click installation and comes integrated with multiple features including image pre-compression, image tagging, and tag statistics. Recently, we also launched the webui plugin version of this tool, everyone is welcome to use it!
This model is a run-accelerated version of the HelloWorld SDXL base model, combining both SDXL Turbo and LCM technologies. Paired with the Eular a sampler, it can generate images within 6-8 steps, which is 3 times faster than the original SDXL version.This model is optimized for the Eular a sampler and it is recommended to only use the Eular a sampler for output.
After multiple rounds of testing, we identified the optimal integration ratio of the SDXL Turbo and SDXL LCM models. The current test results show that for the same 8-step image generation, the effect is: Turbo+LCM dual fusion > Turbo single fusion > LCM single fusion.
The image quality of the 8-step output from the Turbo+LCM dual fusion version is very close to the HelloWorld original model!
The memory usage of the Turbo+LCM dual fusion version is consistent with the HelloWorld original version. Therefore, if you have enough memory, it is recommended to enlarge the direct output image by 1.5 times (still within 6-8 steps).
The recommended parameters for generating images with this model are:
Sampler: Eular a (Important! The model is specifically adapted to Eular a, other samplers may not yield as good results)
CFG scale: 2 (Important! It is recommended to have a CFG scale between 1.5~2.5)
Sampling steps: 8 steps (6~8 steps are acceptable)
Hires algorithm: ESRGAN 4x (Other upscaling algorithms can also be used, not a mandatory option. Please ensure that your GPU memory is sufficient)
Hires Upscale factor: 1.5x
Hires steps: 8 steps
Hires Denoising strength: 0.3