arxiv CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching