"Who's going to take the meeting minutes?" For many professionals, it's an all-too-familiar question. Today, many professionals simply record meetings using Naver Clova Note, where AI automatically converts speech into text, summarizes the key points, and generates meeting minutes. Tasks that once required hours of replaying recordings and manually transcribing conversations can now be completed in just minutes.
But is Speech-to-Text (STT) technology useful only for documenting meetings? Far from it. In healthcare, it is already being used to automatically generate clinical records. In the fashion industry, it is evolving into a productivity tool that can generate designers' fitting feedback and technical specifications in real time.
When discussing the era of generative AI, most people think first of AI-generated advertising images or digital catalogs. Yet the greatest productivity gains in fashion may come not from images, but from speech. Every day, designers, technical designers, and production managers exchange countless verbal instructions and discussions. The moment those conversations are transformed into digital data, repetitive documentation and communication bottlenecks can be eliminated, paving the way for AI-powered workflow automation.
This vision is already becoming reality in healthcare. Oracle and NVIDIA have introduced an AI-powered voice system that transcribes physicians' spoken notes in real time, accurately recognizes medical terminology, and automatically stores the information in Electronic Health Record (EHR) systems. Previously, clinicians had to create EHR documentation after patient consultations using voice memos or handwritten notes. Now, documentation happens simultaneously with patient care.
At the core of this system is NVIDIA Riva, a GPU-accelerated speech AI development platform that improves speech recognition accuracy by learning specialized medical terminology. Running on Oracle Cloud Infrastructure (OCI) Kubernetes environments, it delivers reliable performance.
The significance of this technology extends well beyond basic speech transcription. During speech recognition, NVIDIA Riva removes background noise, accounts for accents and unique pronunciation, and converts spoken language into highly accurate text. Its language models then resolve homonyms by selecting the words and sentences that best fit the surrounding context. Translation models can also be incorporated into the workflow.
The fashion industry is well-positioned to adopt the same approach. Today, design reviews, sample revisions, Tech Pack creation, and communication with manufacturing partners still rely heavily on unstructured channels such as phone calls, messaging platforms, and email. For example, a designer may say, "Please lower the neckline by 1.5 cm," or "Increase the pant hem width by 1 cm." In many cases, a technical designer must manually enter these revisions into the Tech Pack or relay them verbally to others.
If revision requests are omitted or incorrectly documented, mistakes can propagate through the Tech Pack, resulting in costly rework, production delays, and unnecessary sample development — ultimately increasing both lead times and operating costs.
This is precisely where Voice-First AI comes into play. Speech AI trained on fashion-specific terminology can transcribe a designer's spoken instructions in real time while automatically categorizing revisions by pattern, size, and garment component. Vision AI can then analyze photographs of revised garment areas and insert them into the appropriate sections of the Tech Pack. Once the designer completes the final review through a human-in-the-loop process, the updated Tech Pack can be automatically distributed across the PLM system, material teams, pattern makers, production management, and sewing factories.
The ultimate value of this transformation is not reducing headcount — it is eliminating waiting time. The fashion industry's biggest bottlenecks stem from unstructured tacit knowledge, fragmented data, and repetitive documentation. By combining Speech-to-Text technology with AI, these workflows can be automated end to end. As a result, fashion designers are freed to focus on what matters most: creative design work.
In the Voice-First era, speech becomes documentation, documentation becomes data, and the future of fashion innovation may depend not on AI that draws better, but on AI that listens better.
"중요한 회의인데 회의록은 누가 작성하지?" 직장인이라면 한 번쯤 해본 고민이다. 실제로 많은 사람들이 네이버 클로바노트를 이용해 회의를 녹음한 뒤 AI가 음성을 텍스트로 변환하고 핵심 내용을 자동 요약해 회의록을 작성한다. 기존에는 녹음 파일을 다시 들으며 몇 시간 걸리던 녹취가 이제는 몇분 안에 끝난다. 그렇다면 이러한 STT (Speech-to-Text) 기술은 회의록 작성에만 활용될까? 의료 현장에서는 진료 기록을 자동 작성하고, 패션 산업에서는 디자이너의 피팅 피드백과 작업지시서를 실시간으로 생성하는 생산성 혁신으로 진화하고 있다.
생성형 AI 시대를 이야기할 때 많은 사람들은 광고 이미지 생성이나 디지털 카탈로그를 먼저 떠올린다. 그러나 패션 산업에서 가장 큰 생산성 향상은 이미지가 아닌 '음성(Speech-to-Text)'에서 시작될 가능성이 높다. 디자이너, 테크니컬 디자이너, 생산관리자가 매일 주고 받는 수많은 대화가 디지털 데이터로 전환된는 순간, 반복적인 문서작업과 커뮤니케이션 병목이 사라지고 AI에 의한 업무 자동화가 가능해지기 때문이다.
이러한 가능성은 이미 의료 산업에서 현실이 되고 있다. Oracle과 Nvidia는 의료 AI료 음성 시스템을 통해 의사가 말한 내용을 AI가 실시간으로 텍스트화하고 의료 전문 용어까지 정확하게 인식해 전자 의료기록 시스템(EHR)에 자동 저장하는 구조를 선보였다. 기존에는 의료진이 진료 후 음성메모나 수기 기록을 바탕으로 다시 EHR를 작성해야 했지만, 이제는 진료와 기록이 동시에 이뤄진다. 엔비디아 Riva는 GPU 기반 음성 AI 소프트웨어 개발 플랫폼으로 의료 전문용어 학습을 통해 AI 음성 인식의 정확도를 높였으며, Oracle Cloud Infrastructure(OCI)의 쿠버네티스 환경 위에서 안정적으로 서비스를 제공한다.
이 사례의 핵심은 단순한 STT가 아니다. 기술을 설명하면, 엔비디아 Riva는 음성인식 단계에서는 말소리에서 잡음을 없애고 억양, 특이한 발음을 고려해 정확한 텍스트로 잡아내고 language model을 이용해서 동음 이의어 단어 중에서 가장 문맥에 맞는 단어와 문장으로 바꿔준다. 그 이후에 번역 모델도 추가할 수 있다.
패션 산업 역시 이러한 접근이 가능하다. 현재는 디자인 검토, 샘플 수정작업, 작업 지시서(Tech Pack) 작성, 공장과의 커뮤니케이션은 대부분 전화, 메신저, 메일 등 비정형 데이터에 의존한다. 예를 들어 디자이너가 "네크라인을 1.5cm 낮춰주세요." "팬츠 밑단폭을 1cm 늘려 주세요"라고 피팅 수정 의견을 전달하면, 테크니컬 디자이너가 이를 작업지시서에 직접 입력하거나 구두로 전달하는 경우가 많다.
이 과정에서 소통 오류나 실수로 수정 사항이 누락되면 작업지시서 오기재로 이어지고, 결국 봉제 재작업과 납기 지연, 불필요한 샘플 제작 등 막대한 시간과 비용 손실로 이어질 수 있다.
이를 해결하는 것이 바로 'Voice-First AI'이다. 패션 전문용어를 학습한 음성 AI는 디자이너의 음성을 실시간으로 텍스트화하고 수정 내용을 패턴·사이즈·부위별 데이터로 자동 분류한다. 이어 비전AI는 촬영한 수정 부위 사진을 정확한 작업지시서 위치에 삽입한다. 디자이너가 최종 검토(Human-in-the-loop)를 마치면 수정된 작업지시서는 PLM 시스템과 소재실, 패턴실, 생산관리, 봉제 공장까지 자동 배포된다.
이러한 변화의 핵심은 사람을 줄이는 것이 아니라 대기 시간을 없애는 것이다. 패션 산업의 병목은 비정형 암묵지, 데이터 파편화와 반복적인 문서작업인데 STT와 AI는 자동화 업무 프로세스를 가능하게 한다. 이로써 패션 디자이너는 핵심 업무인 '창의적인 기획'에 집중할 수 있다. 음성은 기록이 되고 기록은 데이터가 되며, 앞으로 패션 산업의 혁신은 '잘 그리는 AI'가 아니라 '잘 듣는 AI'일지도 모른다.