Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jbhuang0604
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
How I Understand Flow Matching [video]
(youtube.com)
2 points
by
jbhuang0604
2y ago
|
0 comments
2.
▲
3D human digitization from a single image [video]
(youtube.com)
1 points
by
jbhuang0604
3y ago
|
0 comments
3.
▲
AI 3D Generation [video]
(youtube.com)
2 points
by
jbhuang0604
3y ago
|
0 comments
4.
▲
by
jbhuang0604
3y ago
Thank you!
5.
▲
by
jbhuang0604
3y ago
The UI part is basically a way to organize the user's intention. In the backend, we develop method for extracting "token maps" (i.e., which spatial regions correspond to specific words) and use region-based diffusion to achie
6.
▲
by
jbhuang0604
3y ago
Yes! I am particularly excited about this feature.
7.
▲
by
jbhuang0604
3y ago
Yup, it could be similar, but it mostly only works for very simple prompts (e.g., one subject in the image). For example, in Figure 11 of the paper ( https://arxiv.org/pdf/2304.06720.pdf ), you can see that full-text &qu
8.
▲
by
jbhuang0604
3y ago
Yes, if you go to the huggingface demo: https://huggingface.co/spaces/songweig/rich-text-to-image You can find the segmentation information on the bottom-right of the rich-text generation result.
9.
▲
by
jbhuang0604
3y ago
Got it! Thanks for the feedback! This is definitely something we can improve.
10.
▲
ClimateNeRF: Extreme Weather Synthesis in Neural Radiance Field
(climatenerf.github.io)
1 points
by
jbhuang0604
3y ago
|
0 comments
11.
▲
by
jbhuang0604
3y ago
Thanks for the comment! Our method is model-agnostic . It can be easily adapted to any LLM (aka the text-encode) and any text-to-image models. For example, the method was originally tested in Stable Diffusion 1.4. But we can easily apply i
12.
▲
by
jbhuang0604
3y ago
Yes, I think the "footnote" showcases this well. You can use it to interactively explore your visual imagination. Some examples here: https://youtu.be/ihDbAUh0LXk?si=i3LFfkDXIDKKvne3&t=91
13.
▲
by
jbhuang0604
3y ago
Thanks! BTW, we recently implemented the Automatic1111 extension so you can install and use it directly with A1111. https://github.com/songweige/sd-webui-rich-text
14.
▲
by
jbhuang0604
3y ago
We compared with three other prompt term weighting (aka font size) methods: - Prompt-to-Prompt - Parenthesis - Repeating Our results show that Prompt-to-Prompt would generate artifacts and other heuristic methods are not very effective. You
15.
▲
by
jbhuang0604
3y ago
Yes, the particular font family was just a "label". We don't really interpret what that font style would look like.
16.
▲
by
jbhuang0604
3y ago
Exactly! You get to have full control over the detailed contents you wish to generate.
17.
▲
by
jbhuang0604
3y ago
Yes, you can expand the rich text information into a long sentence. We call this full-text in the paper. The issue of using "full-text" is that it's hard to edit the image interactively. Every time you change the text, you ge
18.
▲
by
jbhuang0604
3y ago
Thanks a lot for the comment! (one of the authors here) RE: plaintext - The "plain-text" result is just a baseline. We call the "plaintext parentheses in the prompt" full-text (i.e., expanding the rich text info into a l
19.
▲
Expressive Text-to-Image Generation with Rich Text [video]
(youtube.com)
1 points
by
jbhuang0604
3y ago
|
0 comments
20.
▲
Consistent Video Depth Estimation
(youtube.com)
2 points
by
jbhuang0604
6y ago
|
1 comments
21.
▲
by
jbhuang0604
6y ago
Yes, you are right. Existing single-image depth estimation models are not far from perfect. We hope to see future development in that direction to further improve the visual quality of 3D photos.
22.
▲
by
jbhuang0604
8y ago
Thanks for sharing. Yes, what we found is that multiple input images often are not consistent with each other (i.e., no physical volume that can satisfy the constraints from all the views). We thus start with a set of consistent voxels and
23.
▲
by
jbhuang0604
8y ago
This is interesting. Imagine that we have a series of multi-view wire structure (i.e., adding a time dimension), then we probably can project three different animations.
24.
▲
Multi-view Wire Art
(cgv.cs.nthu.edu.tw)
73 points
by
jbhuang0604
8y ago
|
10 comments
25.
▲
Multi-view Wire Art
(bit.ly)
1 points
by
jbhuang0604
8y ago
|
1 comments
26.
▲
Awesome Computer Vision
(github.com)
5 points
by
jbhuang0604
12y ago
|
0 comments