Carbon-Silicon Synergy: Architectural AI Design with Fused Text-Image Data

Authors

  • Yuxuan Zhu Gold Mantis School of Architecture, Soochow University, Suzhou 215009, China
  • Shu Wang College of Architecture, Nanjing Tech University, Nanjing 211816, China

DOI:

https://doi.org/10.53469/jpce.2025.07(10).04

Keywords:

Carbon-silicon synergy, Architectural AI, Multimodal design, Text-image data, Human-AI collaboration, Design workflow

Abstract

Architectural design increasingly operates across two interdependent environments: the material setting in which buildings are constructed and inhabited, and the computational setting in which requirements, images, simulations, and design alternatives are produced. This paper uses the term carbon-silicon synergy to describe a working relationship between these environments rather than a progression from the physical to the virtual. Its focus is the role of fused text-image data in architectural AI. Text can express programme, intention, regulation, and qualitative priorities, whereas images and drawings convey geometry, appearance, spatial hierarchy, and context. When these modalities are aligned, they can support a design process in which computational models translate between briefs, visual evidence, and candidate schemes while architects retain responsibility for interpretation and choice. The paper develops a conceptual framework organized around three requirements: complementary representation, traceable translation, and reversible human intervention. It then outlines a human-in-the-loop workflow extending from brief formulation and site interpretation to controlled generation, comparative evaluation, and documentation. Existing techniques in semantic segmentation, vision-language representation, diffusion-based synthesis, and graph-constrained layout generation are discussed as components of this workflow rather than as autonomous design agents. The analysis also identifies limits concerning constructability, cultural specificity, data provenance, authorship, and model error. The proposed framework is methodological; it does not report a trained system or empirical validation. Its contribution is to clarify how multimodal AI may be incorporated into architectural practice without reducing design to image production or transferring professional judgment to an opaque model.

Downloads

Published

2025-10-31

How to Cite

Zhu, Y., & Wang, S. (2025). Carbon-Silicon Synergy: Architectural AI Design with Fused Text-Image Data. Journal of Progress in Civil Engineering, 7(10), 33–37. https://doi.org/10.53469/jpce.2025.07(10).04

Issue

Section

Articles