Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

HwpForge

한글 문서(HWP/HWPX)를 프로그래밍으로 제어하는 Rust 라이브러리

crates.io docs.rs License: MIT OR Apache-2.0


HwpForge란?

HwpForge는 한컴 한글의 HWP/HWPX 문서를 Rust로 읽고, 쓰고, 변환할 수 있는 라이브러리입니다. public guide와 umbrella crate surface는 여전히 HWPX/Markdown 중심이지만, HWP5는 전용 crate와 CLI를 통해 decode, audit, re-emission 경로를 제공합니다.

주요 기능

  • HWPX 풀 코덱 — HWPX 파일 디코딩/인코딩 + 무손실 라운드트립
  • HWP5 읽기/점검 경로 — legacy .hwp decode, audit, HWPX re-emission
  • Markdown 브릿지 — GFM Markdown ↔ HWPX 변환
  • PDF 내보내기 — 문서에 들어 있는 조판 캐시(layout cache)를 재생해 PDF로 렌더 (to-pdf)
  • 편집 연산 (ops) — 누름틀 채우기, 표 셀 편집, 문단 삽입·삭제, 템플릿 스탬핑, diff·검증을 CLI·MCP·Python이 같은 구현으로 공유
  • Python 바인딩 — pip install hwpforge
  • MCP 서버 — AI 도구에서 한글 문서를 직접 생성·편집 (MCP 레퍼런스)
  • YAML 스타일 템플릿 — 재사용 가능한 디자인 토큰 (Figma 패턴)
  • 타입 안전 API — 브랜드 인덱스, 타입스테이트 검증, unsafe 코드 0

지원 콘텐츠

카테고리요소
텍스트런, 문자 모양, 문단 모양, 스타일 (한컴 기본 스타일 18/22/23개, 스타일 세트별)
구조표 (중첩), 이미지, 글상자, 캡션
레이아웃다단, 페이지 설정, 가로/세로, 여백, 마스터페이지
머리글/바닥글머리글, 바닥글, 쪽번호 (autoNum)
주석각주, 미주
도형선, 사각형, 타원, 다각형, 호, 곡선, 연결선, 묶음 객체, 글맵시 (채움/회전/화살표)
글자 효과덧말, 글자겹침
수식HancomEQN 스크립트
차트18종 차트 (OOXML 호환)
참조책갈피, 상호참조, 필드, 메모, 찾아보기
MarkdownGFM 디코드, 손실/무손실 인코드, YAML 프론트매터

누구를 위한 라이브러리인가?

  • LLM/AI 에이전트 — 자연어로 한글 문서 자동 생성
  • 백엔드 개발자 — 서버에서 한글 문서 프로그래밍 생성
  • 자동화 도구 — CI/CD 파이프라인에서 보고서 자동 생성
  • 데이터 파이프라인 — HWPX 문서에서 텍스트/표 추출

빠른 맛보기

#![allow(unused)]
fn main() {
use hwpforge::core::{Document, Draft, Paragraph, Run, Section, PageSettings};
use hwpforge::foundation::{CharShapeIndex, ParaShapeIndex};
use hwpforge::hwpx::{HwpxEncoder, HwpxStyleStore};
use hwpforge::core::ImageStore;

// 1. 문서 생성
let mut doc = Document::<Draft>::new();
doc.add_section(Section::with_paragraphs(
    vec![Paragraph::with_runs(
        vec![Run::text("안녕하세요, HwpForge!", CharShapeIndex::new(0))],
        ParaShapeIndex::new(0),
    )],
    PageSettings::a4(),
));

// 2. 검증 + 인코딩
let validated = doc.validate().unwrap();
let style_store = HwpxStyleStore::with_default_fonts("함초롬바탕");
let image_store = ImageStore::new();
let bytes = HwpxEncoder::encode(&validated, &style_store, &image_store).unwrap();

// 3. 파일 저장
std::fs::write("output.hwpx", &bytes).unwrap();
}

다음 단계

설치

HwpForge는 순수 Rust로 작성된 라이브러리입니다. 별도의 시스템 의존성 없이 Cargo.toml에 추가하는 것만으로 사용할 수 있습니다.

최소 지원 Rust 버전 (MSRV)

라이브러리 크레이트(hwpforge umbrella와 하위 크레이트 대부분)는 Rust 1.89 이상이 필요합니다. 현재 버전을 확인하려면:

rustc --version

버전이 낮다면 rustup으로 업데이트합니다:

rustup update stable

PDF 렌더(krilla)에 의존하는 hwpforge-smithy-pdf, hwpforge-convert, hwpforge-bindings-cli, hwpforge-bindings-py는 각 Cargo.toml에 rust-version = "1.92"가 지정돼 있어 Rust 1.92 이상이 필요합니다. hwpforge umbrella만 쓰는 라이브러리 사용자는 1.89면 충분하고, CLI를 소스에서 설치하거나 Python sdist를 빌드할 때만 1.92가 필요합니다.

의존성 추가

Cargo.toml의 [dependencies] 섹션에 추가합니다:

[dependencies]
hwpforge = "0.16"

기본 설치에는 HWPX 인코더/디코더가 포함됩니다.

Feature Flags

HwpForge는 필요한 기능만 선택적으로 활성화할 수 있습니다.

Feature기본 포함설명
hwpx예HWPX 인코더/디코더 (ZIP + XML, KS X 6101)
md아니오Markdown(GFM) ↔ HWPX 변환
ops-hwpx아니오HWPX 전용 표면(검사·교환·읽기·편집·diff·stamp·스타일)이 공유하는 연산 계층. hwpx를 함께 켭니다
ops-md아니오연산 계층에 Markdown 연산(convert_md, to_md)을 더합니다. ops-hwpx와 md를 함께 켭니다
ops아니오연산 계층 전체의 별칭(ops-md)
schemars아니오ops가 소유한 wire DTO와 교환 DTO에 JsonSchema derive를 더합니다. HWPX 코덱을 따라 켜지지는 않습니다
full아니오hwpx + md + ops를 켭니다(schemars는 포함하지 않음)

HWPX만 사용 (기본)

[dependencies]
hwpforge = "0.16"

Markdown 변환 포함

[dependencies]
hwpforge = { version = "0.16", features = ["md"] }

편집 연산 계층 포함

CLI·MCP·Python이 공유하는 연산(hwpforge::ops)을 라이브러리에서 직접 쓰려면 ops를 켭니다.

[dependencies]
hwpforge = { version = "0.16", features = ["ops"] }

모든 기능 활성화

[dependencies]
hwpforge = { version = "0.16", features = ["full"] }

빌드 확인

의존성을 추가한 후 빌드가 정상적으로 되는지 확인합니다:

cargo build

다음과 같이 컴파일이 성공하면 설치가 완료된 것입니다:

Compiling hwpforge v0.16.9
 Finished `dev` profile [unoptimized + debuginfo] target(s) in ...

CLI · MCP · Python

Rust 라이브러리 외에 다른 창구도 있습니다. 설치 방법은 각 장에 있습니다.

  • CLI (hwpforge): crates.io에 배포하지 않으므로 git 또는 로컬 경로에서 설치합니다. Rust 1.92 이상이 필요합니다. CLI 레퍼런스
  • MCP 서버 (hwpforge-mcp): npx -y @hwpforge/mcp 또는 cargo install hwpforge-bindings-mcp. MCP 서버 레퍼런스
  • Python: pip install hwpforge. Python 가이드

다음 단계

설치가 완료되었습니다. 빠른 시작으로 이동하여 첫 번째 HWPX 문서를 생성해 보세요.

빠른 시작

이 페이지에서는 HwpForge의 세 가지 핵심 사용 패턴을 코드 예제와 함께 설명합니다.

세 예제 모두 hwpforge 하나만 의존성으로 두면 그대로 컴파일됩니다 — 오류 타입은 표준 라이브러리의 Box<dyn std::error::Error>를 쓰므로 오류 처리 크레이트를 따로 추가할 필요가 없습니다. 예제 3은 features = ["md"]가 필요합니다(설치).

예제 1: 텍스트 문서 생성 후 HWPX로 저장

가장 기본적인 사용 패턴입니다. 문서 구조를 직접 조립하고 HWPX 파일로 출력합니다.

use hwpforge::core::{Document, Draft, ImageStore, PageSettings, Paragraph, Run, Section};
use hwpforge::foundation::{CharShapeIndex, ParaShapeIndex};
use hwpforge::hwpx::{HwpxEncoder, HwpxStyleStore};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // 1. Draft 상태의 문서 생성
    let mut doc = Document::<Draft>::new();

    // 2. 텍스트 Run 구성 — CharShapeIndex(0)은 기본 글자 스타일을 참조
    let run = Run::text("안녕하세요, HwpForge입니다!", CharShapeIndex::new(0));

    // 3. 문단 생성 — ParaShapeIndex(0)은 기본 문단 스타일을 참조
    let paragraph = Paragraph::with_runs(vec![run], ParaShapeIndex::new(0));

    // 4. 섹션(= 쪽 단위 컨테이너)에 문단 추가, A4 용지 설정
    let section = Section::with_paragraphs(vec![paragraph], PageSettings::a4());
    doc.add_section(section);

    // 5. 유효성 검증 — Draft → Validated 상태 전이 (타입스테이트)
    let validated = doc.validate()?;

    // 6. 스타일 스토어: 한컴 Modern(22종) 기본 스타일 사용
    let style_store = HwpxStyleStore::with_default_fonts("함초롬바탕");

    // 7. 이미지 스토어: 이미지가 없으므로 빈 스토어 사용
    let image_store = ImageStore::new();

    // 8. HWPX 바이트 인코딩 후 파일 저장
    let bytes = HwpxEncoder::encode(&validated, &style_store, &image_store)?;
    std::fs::write("output.hwpx", &bytes)?;

    println!("output.hwpx 저장 완료 ({} bytes)", bytes.len());
    Ok(())
}

참고: CharShapeIndex::new(0)과 ParaShapeIndex::new(0)은 HwpxStyleStore::with_default_fonts()이 제공하는 기본 스타일(본문)을 가리킵니다. 커스텀 스타일을 사용하려면 스타일 템플릿 문서를 참고하세요.


예제 2: HWPX 파일 디코딩

기존 HWPX 파일을 읽어서 Core 문서 모델로 변환합니다.

use hwpforge::hwpx::HwpxDecoder;
use hwpforge::core::run::RunContent;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // 파일 경로를 받아 HWPX를 디코딩
    let result = HwpxDecoder::decode_file("input.hwpx")?;

    let doc = &result.document;

    // 섹션 수 출력
    println!("섹션 수: {}", doc.sections().len());

    // 메타데이터 접근 (제목, 작성자, 작성일 등)
    let meta = doc.metadata();
    if let Some(title) = &meta.title {
        println!("제목: {}", title);
    }
    if let Some(author) = &meta.author {
        println!("작성자: {}", author);
    }

    // 각 섹션의 문단과 텍스트 출력
    for (sec_idx, section) in doc.sections().iter().enumerate() {
        println!("--- 섹션 {} ---", sec_idx + 1);
        for paragraph in &section.paragraphs {
            for run in &paragraph.runs {
                if let RunContent::Text(ref text) = run.content {
                    print!("{}", text);
                }
            }
            println!(); // 문단 끝 줄바꿈
        }
    }

    Ok(())
}

HwpxDecoder::decode_file은 경로를 받아 ZIP을 열고 XML을 파싱합니다. 반환값에는 document(문서 구조), style_store(글꼴/문단 스타일), image_store(이미지), warnings(디코드 중 표면화된 비치명 경고)가 포함됩니다. document.metadata()로 제목, 작성자 등의 메타데이터에 접근할 수 있습니다.


예제 3: Markdown → HWPX 변환

GFM(GitHub Flavored Markdown) 텍스트를 HWPX 파일로 변환합니다. features = ["md"] 또는 features = ["full"]이 필요합니다.

use hwpforge::core::ImageStore;
use hwpforge::hwpx::{HwpxEncoder, HwpxRegistryBridge};
use hwpforge::md::MdDecoder;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    // 1. GFM Markdown 텍스트 (YAML 프론트매터 지원)
    let markdown = r#"---
title: 보고서 제목
author: 홍길동
date: 2026-03-06
---

1장. 서론

본 보고서는 HwpForge를 이용한 **자동 문서 생성** 예시입니다.

# 1.1 배경

- 항목 A
- 항목 B
- 항목 C

# 1.2 결론

> HwpForge는 LLM 에이전트가 한글 문서를 생성할 때 사용할 수 있습니다.
"#;

    // 2. Markdown → Core 문서 모델 변환
    //    MdDecoder::decode_with_default는 document + style_registry(스타일 매핑)를 반환
    let md_doc = MdDecoder::decode_with_default(markdown)?;

    // 3. registry-local 스타일 인덱스를 HWPX store-local 인덱스로 rebinding
    let bridge = HwpxRegistryBridge::from_registry(&md_doc.style_registry)?;
    let rebound = bridge.rebind_draft_document(md_doc.document)?;

    // 4. Draft → Validated 상태 전이
    let validated = rebound.validate()?;

    // 5. 이미지 없음
    let image_store = ImageStore::new();

    // 6. HWPX 인코딩 후 저장
    let bytes = HwpxEncoder::encode(&validated, bridge.style_store(), &image_store)?;
    std::fs::write("report.hwpx", &bytes)?;

    println!("report.hwpx 저장 완료");
    Ok(())
}

Markdown 변환 시 자동으로 처리되는 항목:

Markdown 요소변환 결과
# H1 ~ ###### H6한컴 개요 1 ~ 6 스타일
**굵게**글자 진하게
*기울임*글자 기울임
`코드`고정폭 글꼴
> 인용문들여쓰기 문단
- 목록글머리 기호 목록
YAML 프론트매터문서 메타데이터

다음 단계

아키텍처 개요

HwpForge는 대장간(Forge) 메타포를 기반으로 설계된 계층형 크레이트 구조를 갖습니다. 각 계층은 명확한 역할을 가지며, 상위 계층은 하위 계층에만 의존합니다.

Forge 메타포

계층역할비유
Foundation (기반)원시 타입, 단위, 인덱스쇠못과 금속 소재
Core (핵심)형식 독립 문서 모델도면 위의 설계도
Blueprint (청사진)YAML 스타일 템플릿피그마 디자인 토큰
Smithy (대장간)형식별 인코더/디코더 (PDF는 쓰기 전용 렌더러)용광로와 망치
Convert (변환)포맷 간 변환 오케스트레이터단조 작업 지휘
Bindings (바인딩)Python, CLI, MCP 인터페이스완성된 제품 포장

크레이트 의존성 그래프

graph TD
    F[hwpforge-foundation<br/>원시 타입] --> C[hwpforge-core<br/>문서 모델]
    F --> B[hwpforge-blueprint<br/>스타일 템플릿]
    C --> B
    F --> SH[hwpforge-smithy-hwpx<br/>HWPX 코덱]
    C --> SH
    B --> SH
    F --> SM[hwpforge-smithy-md<br/>Markdown 코덱]
    C --> SM
    B --> SM
    F --> S5[hwpforge-smithy-hwp5<br/>HWP5 decode/projection]
    C --> S5
    F --> SPDF[hwpforge-smithy-pdf<br/>레이아웃 캐시 재생 렌더러]
    C --> SPDF
    F --> CONV["hwpforge-convert<br/>HWP5 → HWPX 오케스트레이터"]
    C --> CONV
    SH --> CONV
    S5 --> CONV
    SPDF --> CONV
    F --> U[hwpforge<br/>umbrella crate]
    C --> U
    B --> U
    SH --> U
    SM --> U
    F --> CLI["hwpforge-bindings-cli<br/>CLI (shipped)"]
    C --> CLI
    U --> CLI
    SH --> CLI
    SM --> CLI
    S5 --> CLI
    SPDF --> CLI
    CONV --> CLI
    F --> MCP["hwpforge-bindings-mcp<br/>MCP (shipped)"]
    C --> MCP
    U --> MCP
    SH --> MCP
    SM --> MCP
    F --> PY["hwpforge-bindings-py<br/>Python"]
    U --> PY
    CONV --> PY
    SPDF --> PY

화살표는 Cargo.toml의 [dependencies] 직접 의존을 나타내며(A --> B는 B가 A에 의존), 개발용 의존성([dev-dependencies])은 제외했습니다. umbrella의 hwpforge-smithy-hwpx·hwpforge-smithy-md 의존은 각각 feature hwpx·md가 켜질 때만 생깁니다.

규칙: 의존성은 위에서 아래로만 흐릅니다. foundation을 수정하면 모든 크레이트가 재빌드됩니다. 따라서 foundation은 최소한으로 유지합니다.

공유 연산 계층: hwpforge::ops

공유 문서 연산은 umbrella crate hwpforge의 ops 모듈(feature ops)에 있고, HWP5 변환·PDF 렌더링 연산은 hwpforge-convert의 ops 모듈에 있습니다. 세 바인딩(CLI·MCP·Python)은 각자 포맷을 다루는 대신 이 연산 계층을 호출하며, 에러 코드 표(OpsCode)는 하나로 정의됩니다. 기존 공개 계약이 있는 CLI와 MCP는 레거시 코드 문자열을 각자의 호환 테이블로 유지하고, Python은 OpsCode 문자열을 직접 노출합니다.

핵심 원칙: 구조와 스타일의 분리

HwpForge는 HTML + CSS의 관계처럼 문서 구조와 스타일 정의를 완전히 분리합니다.

Core (구조)           Blueprint (스타일)
─────────────         ──────────────────
Paragraph             font: "맑은 고딕"
  style_id: 2    ──▶  size: 10pt
  runs: [...]         color: #000000
  • Core는 스타일 ID(인덱스)만 보유합니다. 실제 글꼴 이름이나 크기를 모릅니다.
  • Blueprint는 스타일 정의를 YAML 템플릿으로 관리합니다.
  • Smithy 컴파일러가 Core + Blueprint를 조합해 최종 형식을 생성합니다.

이 구조 덕분에 하나의 YAML 템플릿을 여러 문서에 재사용하거나, 동일한 문서를 HWPX/Markdown 등 다른 형식으로 내보낼 수 있습니다.

타입스테이트 패턴: Document → Document

Document는 컴파일 타임에 상태를 추적하는 타입스테이트 패턴을 사용합니다.

#![allow(unused)]
fn main() {
use hwpforge::core::{Document, Draft, PageSettings, Paragraph, Run, Section};
use hwpforge::foundation::{CharShapeIndex, ParaShapeIndex};

// Draft 상태: 편집 가능, 저장 불가
let mut doc = Document::<Draft>::new();
doc.add_section(Section::with_paragraphs(
    vec![Paragraph::with_runs(
        vec![Run::text("본문", CharShapeIndex::new(0))],
        ParaShapeIndex::new(0),
    )],
    PageSettings::a4(),
));

// validate()를 호출해야만 Validated 상태로 전이
let validated = doc.validate().unwrap();

// Validated 상태에서만 인코딩 가능
// doc.validate()를 건너뛰면 컴파일 에러 발생
let bytes = hwpforge::hwpx::HwpxEncoder::encode(
    &validated,
    &hwpforge::hwpx::HwpxStyleStore::with_default_fonts("함초롬바탕"),
    &hwpforge::core::ImageStore::new(),
).unwrap();
}

잘못된 상태에서 저장을 시도하면 런타임 에러가 아닌 컴파일 에러가 발생합니다.

이중 포맷 설계: HWP5 + HWPX

한국에는 두 가지 주요 문서 포맷이 있습니다:

  • HWP5 (.hwp): OLE2/CFB 바이너리 컨테이너 + TLV 레코드 (1990년대~현재, 레거시)
  • HWPX (.hwpx): ZIP 컨테이너 + XML 파일 (KS X 6101 국가 표준, 2014년~현재)

HwpForge는 Core DOM이 포맷에 독립적이도록 설계하여 두 포맷을 통합 처리합니다:

HWP5 (.hwp)  ──decode──▶ ┌────────────────────┐ ◀──decode── Markdown (.md)
                          │  Document<Draft>   │
HWPX (.hwpx) ──decode──▶ │  (포맷 독립 IR)    │ ──encode──▶ HWPX / Markdown
                          └────────────────────┘

모든 Smithy 크레이트는 Core DOM으로/에서 변환만 수행합니다. 비즈니스 로직은 Core에만 의존하므로, 새 포맷(예: smithy-odt)을 추가해도 기존 코드를 수정할 필요가 없습니다.

현재 HWP5 경로는 hwpforge-smithy-hwp5와 CLI surface에서 실사용 가능하며, umbrella crate 중심 예제는 여전히 HWPX/Markdown 쪽이 먼저 소개됩니다.

자세한 내용은 HWP5와 HWPX: 이중 포맷 파이프라인을 참고하세요.

각 크레이트 설명

hwpforge-foundation

의존성이 없는 루트 크레이트입니다. 모든 크레이트가 공유하는 원시 타입을 정의합니다.

  • HwpUnit: 정수 기반 HWP 단위 (1pt = 100 HWPUNIT). 부동소수점 오차 없음
  • Color: BGR 바이트 순서 색상 타입. Color::from_rgb(r, g, b)로 생성
  • Index<T>: 팬텀 타입을 이용한 브랜드 인덱스. CharShapeIndex와 ParaShapeIndex를 혼용하면 컴파일 에러

hwpforge-core

형식에 독립적인 문서 모델입니다. 한글/Markdown/PDF 어디에도 종속되지 않습니다.

  • Document<S>, Section, Paragraph, Run — 기본 문서 구조
  • Table, Control, Shape — 복합 요소
  • PageSettings — 용지 크기, 여백, 가로/세로 방향

hwpforge-blueprint

YAML로 작성하는 스타일 템플릿 시스템입니다. 피그마의 디자인 토큰 개념과 유사합니다.

  • 상속(extends)과 병합(merge)을 지원하는 PartialCharShape / CharShape 두 타입 구조
  • StyleRegistry — 파싱 후 인덱스를 할당한 최종 스타일 집합

hwpforge-smithy-hwpx

HWPX ↔ Core 변환을 담당하는 핵심 코덱입니다. KS X 6101(OWPML) 국가 표준을 구현합니다.

  • HwpxDecoder — ZIP + XML 파싱 → Core 문서 모델
  • HwpxEncoder — Core 문서 모델 → ZIP + XML 바이트
  • HwpxStyleStore — 한컴 기본 스타일 22종(Modern) 내장

hwpforge-smithy-md

GFM(GitHub Flavored Markdown) ↔ Core 변환을 담당합니다.

  • MdDecoder — Markdown + YAML 프론트매터 → Core
  • MdEncoder — Core → Markdown (손실/무손실 모드)

hwpforge (umbrella crate)

모든 공개 크레이트를 재내보내기(re-export)하는 진입점 크레이트입니다. 사용자는 이 크레이트 하나만 의존성에 추가하면 됩니다.

다음 단계

Python

hwpforge 패키지는 HwpForge Rust 라이브러리를 얇게 감싼 Python 바인딩입니다. HWPX 문서를 읽고, 검사하고, 편집하고, Markdown·JSON·PDF로 내보내며, 옛 HWP5(.hwp)를 HWPX로 변환합니다. 연산의 의미는 CLI·MCP 서버와 같은 연산 계층(hwpforge::ops)에서 오므로 세 창구가 같은 문서에 같은 결과를 내고, 메서드 이름·인자 표기·일부 실패 코드만 Python의 공개 계약으로 따로 정해져 있습니다.

이 페이지는 설치와 핵심 개념을 다루고, 세부는 하위 페이지로 나뉩니다.

페이지내용
문서 읽기와 검사open·inspect·outline·fields·read·validate·diff·stamp_plan
편집fill·patch·set_cell·insert_para/delete_para·stamp·restyle, 한컴 저장 문서에서 되는 것과 안 되는 것
변환과 내보내기convert_md·from_json·convert_hwp5·to_md·to_json·to_pdf
결과·경고·오류결과 객체, 보고서의 warnings, HwpForgeError의 다섯 속성, 코드 표
레시피양식 채우기 파이프라인, 표를 CSV로, 배치 처리, HWP → PDF, pip 없는 호스트
API 요약모든 메서드·함수의 시그니처와 반환형 한 표

설치

pip install hwpforge

uv를 쓴다면 uv add hwpforge(프로젝트) 또는 uv pip install hwpforge(환경)를 사용하세요.

정확 핀(==X.Y.Z)보다 ~=X.Y.Z를 권장합니다. Python 전용 수정은 X.Y.Z.N 형태로 나가는데, 정확 핀은 그 수정을 받지 못합니다.

CPython 3.9+를 커버하는 abi3 wheel을 다섯 플랫폼(Linux manylinux_2_28 x86_64/aarch64, macOS 11+ arm64, macOS 10.12+ x86_64, Windows x64)에 배포하며, sdist(소스 배포)는 Rust 1.92+와 maturin이 필요합니다. 런타임 의존성은 0개입니다. 자유 스레드(free-threaded) 빌드는 stable ABI가 다루지 않아 지원하지 않습니다.

pip 없는 호스트

wheel은 zip 파일이고 런타임 의존성이 없으므로, pip 없이 풀어서 바로 import할 수 있습니다. 기준은 배포판 버전이 아니라 glibc 2.28 이상입니다 — CPython 3.9+와 glibc 2.28+를 갖춘 Linux x86_64 호스트라면(예: Debian 10/11/12, Ubuntu 20.04 이상; 아래 파일명은 x86_64 wheel):

python3 -c "import zipfile; zipfile.ZipFile('hwpforge-<version>-cp39-abi3-manylinux_2_28_x86_64.whl').extractall('/opt/hf')"
PYTHONPATH=/opt/hf python3 -c "import hwpforge; print(hwpforge.__version__)"

wheel 파일 자체는 PyPI의 files 페이지에서 받을 수 있습니다 — pip이 설치할 때 받는 것과 같은 바이트·sha256입니다. GitHub Release는 발행 후 에셋을 붙일 수 없는 불변 객체라, 별도로 첨부되지 않습니다. 오프라인 배포의 전체 절차는 레시피에 있습니다.

30초 예제

import hwpforge

doc = hwpforge.Document.open("form.hwpx")
for field in doc.fields()["fields"]:
    print(field["name"], "→", field["current"])

result = doc.fill({"user_email": "kim@example.com"})
for warning in result.report["warnings"]:
    print(warning["code"], warning["message"])
result.document.save("form-filled.hwpx")

doc은 그대로이고, 채워진 문서는 result.document입니다. 이 재대입 구조가 이 패키지의 전부라고 해도 지나치지 않습니다 — 아래 네 개념이 그 이유입니다.

핵심 개념 네 가지

불변 값 Document

Document는 불변 값입니다. HWPX 패키지의 바이트를 들고 있는 값 객체로, 두 문서는 바이트가 같으면 같고(==, hash), 어떤 메서드도 받은 문서를 바꾸지 않습니다. 속성 대입은 AttributeError입니다. 편집 메서드는 새 문서를 돌려주므로, 결과를 변수에 다시 받아야 합니다.

값과 보고서를 함께 담는 결과

연산은 값과 보고서를 함께 돌려줍니다. 편집은 DocumentResult(document, report), 텍스트 내보내기는 TextResult(text, report), 바이트 내보내기는 BytesResult(data, report)입니다. 셋 다 frozen dataclass라 보고서를 따로 조회할 필요도, 잃어버릴 일도 없습니다. 검사 연산(inspect·outline 등)은 보고서(dict)만 돌려줍니다.

항상 있는 warnings

보고서의 warnings는 항상 있고, 저장 전에 봐야 합니다. 실패하지 않은 연산도 조용히 성공하지 않습니다. templates()·schema()를 뺀 모든 보고서에 warnings: [{code, message, hint?}]가 있으며, 비어 있을 수는 있어도 빠지지는 않습니다. 문서가 완전히 보존되지 않은 곳(예: 재인코드가 조판 캐시를 버림)을 여기서 알려줍니다.

하나뿐인 예외 HwpForgeError

실패는 HwpForgeError 하나이고, code로 분기합니다. 입력을 받아들인 뒤 연산이 거부하면 HwpForgeError를 던지며 code·message·hint·cause·details 다섯 속성을 가집니다. 인자 형이 틀리면 TypeError/ValueError, 파일 I/O는 OSError입니다. 예외 메시지 문자열이 아니라 exc.code 문자열로 분기하세요.

try:
    doc = doc.set_cell(table=0, at="0,1", text="값").document
except hwpforge.HwpForgeError as exc:
    print(exc.code)      # 예: "INPUT_ENTRIES_NOT_CARRIED"
    print(exc.hint)      # 연산이 아는 되는 길, 없으면 None

자주 나오는 용어

하위 페이지는 아래 용어를 설명 없이 씁니다.

용어뜻
누름틀한컴 문서의 입력 칸. 이름과 안내문을 가지며, fill이 이름으로 값을 채웁니다
조판 캐시한컴이 저장할 때 계산해 둔 줄 나눔·줄 위치. PDF 렌더는 이것을 재생하고, 한컴은 다시 저장할 때 새로 계산합니다
보존 우선(preserve-first) 편집바꿀 XML 조각만 고치고 패키지의 나머지 바이트는 그대로 두는 편집(fill·patch)
재인코드문서를 모델로 디코드해 고친 뒤 패키지를 새로 쓰는 것. 모델이 담지 않는 것(조판 캐시, 모르는 엔트리)은 사라집니다. 무엇을 경고하는지는 연산마다 다릅니다
fail-closed결과가 손상될 것 같으면 손상된 결과를 내는 대신 거부하는 정책

입력 크기 상한

Document.open은 파일을 읽을 때 CLI·MCP와 같은 상한(hwpforge._hwpforge.MAX_FILE_SIZE, 100 MB)을 적용하고, 넘으면 HwpForgeError(INPUT_TOO_LARGE)를 던집니다. 상한은 파일이 보고하는 크기가 아니라 읽은 바이트 수에 걸리므로, stat() 크기가 0으로 보이는 FIFO나 프로세스 치환으로도 더 큰 입력을 밀어 넣을 수 없습니다. Document.from_bytes에는 상한이 없습니다 — 호출자가 이미 바이트를 들고 있어 더 제한할 읽기가 남아 있지 않기 때문입니다. 100 MB보다 큰 문서를 다뤄야 한다면 직접 읽어서 from_bytes로 넘기면 됩니다.

읽기와 쓰기 포맷

읽기는 .hwpx, 그리고 hwpforge.convert_hwp5를 거친 .hwp(HWP5)입니다. 변환은 HWPX로 가는 한 방향이라 원본 .hwp로 되돌아가는 길은 없습니다. 쓰기는 .hwpx뿐이며(save는 무엇을 읽었든 HWPX 패키지를 씁니다), 한컴 오피스는 .hwpx를 그대로 엽니다.

예제에 쓰인 파일

하위 페이지의 예제는 저장소의 테스트 픽스처를 다음 이름으로 복사해 둔 것을 전제합니다. 같은 이름으로 준비하면 모든 코드 블록이 적힌 그대로 실행됩니다.

예제 파일원본 (tests/fixtures/…)특징
form.hwpxfields/clickhere_named.hwpx누름틀 user_email 하나
grid.hwpxtables/merged_grid_form.hwpx3×2 표(라벨 성명·비고), HwpForge 생성본
template.hwpxstamp/placeholder_basic.hwpx스탬프 후보 ( )·□
plain.hwpxstructural/plain_paragraphs.hwpx문단 4개, HwpForge 생성본
hancom.hwpxtables/table_01_basic_2x2.hwpx한컴 오피스가 저장한 문서
old.hwpstructural/plain_inserted.hwpHWP5 바이너리

다음 단계

문서 읽기와 검사

읽기 연산은 문서를 바꾸지 않고 보고서(dict)를 돌려줍니다. 문서는 열 때 디코드되지 않습니다 — Document.open은 바이트만 들고, 각 연산이 필요할 때 디코드합니다. 그래서 잘못된 바이트는 open이 아니라 첫 연산에서 HwpForgeError(DECODE_FAILED)로 드러납니다.

열기

import hwpforge

doc = hwpforge.Document.open("hancom.hwpx")           # 파일에서, 100 MB 상한
same = hwpforge.Document.from_bytes(doc.to_bytes())    # 이미 든 바이트에서, 상한 없음

print(repr(doc), len(doc), doc == same)                # Document(<n> bytes) <n> True

Document(data) 생성자도 있지만 open/from_bytes가 바이트의 출처를 말해 주므로 그쪽을 쓰세요. bytes가 아닌 것을 넘기면 TypeError입니다.

inspect() — 무엇이 얼마나 있는가

report = doc.inspect()
print(report["sections"], report["paragraphs"], report["tables"], report["images"], report["charts"])
print(report["metadata"]["title"], report["metadata"]["author"], report["metadata"]["modified"])
print(report["fields"])                     # 누름틀 이름 목록
section = report["section_details"][0]
print(section["top_level_paragraphs"], section["all_tables"], section["has_header"])

metadata는 title·author·subject·description·last_saved_by·created·modified·keywords를 가지며, 날짜 둘은 없으면 None입니다.

styles=True를 주면 문서가 정의한 글꼴·글자 모양·문단 모양 요약이 styles 키로 더해집니다(기본은 생략).

styles = doc.inspect(styles=True)["styles"]
print(styles["fonts"][0])          # {'id': 0, 'face_name': '함초롬돋움', 'lang': 'HANGUL'}
print(styles["char_shapes"][0])    # id, font_id, size_pt, bold, italic, color
print(styles["para_shapes"][0])    # id, alignment, line_spacing

섹션 계수 — 접두사가 곧 범위

section_details[i]의 키는 접두사로 무엇을 셌는지를 말합니다. 네 집합은 서로 포함 관계가 아닙니다 — 접두사마다 재귀해 들어가는 곳이 달라, 어느 행도 다른 행을 넓힌 판본이 아닙니다.

접두사재귀해 들어가는 곳master page캡션
top_level_없음 — 섹션 본문 흐름 자체만아니오아니오
(접두사 없음)셀·글상자·각주/미주·메모·머리말/꼬리말예아니오
deep_머리말/꼬리말·셀·글상자/주석/도형/메모아니오아니오
all_섹션 XML이 담은 전부(묶음 객체의 자식 포함)아니오예

top_level_과 접두사 없는 키는 문단·표·이미지·차트를 세고, deep_은 문단만, all_은 여섯 가지 객체(all_tables·all_images·all_text_boxes·all_lines·all_rectangles·all_polygons)를 셉니다. non_empty_ 중위사는 세는 집합을 바꾸지 않고 “보이는 텍스트가 있는 문단“으로만 좁히므로, top_level_non_empty_paragraphs는 top_level_paragraphs를 넘지 않습니다. all_charts는 없습니다 — 디코더가 섹션 최상위 문단 밖의 차트를 복원하지 못해, 그런 키를 두면 정작 필요한 문서에서 조용히 과소 계수하게 됩니다.

all_ 키는 디코드된 객체를 세지, 원시 XML 요소를 세지 않습니다. CLI --json에는 같은 범위를 예전 철자(tables·images·text_boxes 등)로 내보내는 키가 있는데 그쪽은 원시 스캔이라, 디코더가 표현하지 못하는 요소를 받아들인 문서에서는 더 큰 값이 나올 수 있습니다. 두 값은 교환 가능하지 않습니다. 문단 키에는 이 단서가 붙지 않습니다.

outline() — 제목·표·누름틀·책갈피의 위치

outline = doc.outline()["outline"]
print(outline["title"])                                   # 없으면 None
for h in outline["headings"]:
    print(h["level"], h["text"], h["at"])                 # at = {'section': s, 'para': p}
for t in outline["tables"]:
    print(t["ordinal"], t["rows"], t["cols"], t["addressable"], t["caption"])
print(outline["fields"], outline["bookmarks"])

tables[i]["ordinal"]이 read(table=…)·set_cell(table=…)이 받는 표 번호입니다(0부터). addressable이 거짓인 표는 격자 주소로 셀을 지정할 수 없습니다(병합이 격자를 깨는 경우).

fields() — 누름틀

for f in hwpforge.Document.open("form.hwpx").fields()["fields"]:
    print(f["name"], f["hint"], f["current"], f["section"], f["fillable"])
# user_email 회사 이메일을 입력하세요 회사 이메일을 입력하세요 0 True

fillable이 거짓인 필드는 fill이 FIELD_NOT_FILLABLE로 거부합니다. name은 이름이 없는 필드에서 None 일 수 있습니다.

read() — 한 부분만

문단 범위·표·누름틀 셋 중 정확히 하나를 지정합니다. 보고서의 세 키(paragraphs·table·fields)는 항상 있고, 묻지 않은 것은 None입니다.

plain = hwpforge.Document.open("plain.hwpx")

view = plain.read(section=0, paras="0..2")["paragraphs"]     # 양끝 포함: 0, 1, 2
for p in view["paragraphs"]:
    print(p["at"]["para"], p["kind"], p["text"])

table = hwpforge.Document.open("grid.hwpx").read(table=0)["table"]
for cell in table["cells"]:
    print(cell["row"], cell["col"], cell["row_span"], cell["col_span"], repr(cell["text"]))

fields = hwpforge.Document.open("form.hwpx").read(field="user_email")["fields"]
print(fields[0]["current"])
  • paras는 "시작..끝"이고 양끝 포함입니다. 범위가 섹션을 넘으면 READ_PARA_RANGE_INVALID, 형식이 틀리면 READ_PARAS_INVALID.
  • 문단은 kind로 구분됩니다: "body", "heading"(level 포함), "list"(numbered·level·checked 포함 — checked는 체크 목록이 아니면 None). 문단이 표·그림 등을 품으면 contains 키가 더해집니다.
  • 표는 rows·cols와 셀 목록입니다. 병합된 영역은 기준 셀 하나로만 나오므로 셀 수가 rows × cols보다 적을 수 있습니다.
  • 아무것도 안 주거나 둘 이상 주면 READ_TARGET_REQUIRED, 표 번호가 넘치면 READ_TABLE_OUT_OF_RANGE.

validate() — 두 종류의 실패

report = doc.validate()
print(report["ok"], report["sections"], report["paragraphs"])
for problem in report["errors"]:
    print(problem["code"], problem["message"])

문서가 디코드는 되지만 모델 불변조건을 어기면 예외 없이 ok가 거짓인 보고서를 돌려주고, 이유는 errors에 담깁니다. 반면 디코드 자체가 안 되는 입력은 연산 실패이므로 HwpForgeError(DECODE_FAILED)를 던집니다. 즉 “유효하지 않은 문서“는 반환값으로, “문서가 아닌 바이트“는 예외로 옵니다.

try:
    hwpforge.Document.from_bytes(b"not a document").validate()
except hwpforge.HwpForgeError as exc:
    print(exc.code)        # DECODE_FAILED

CLI·MCP는 같은 필드를 valid라 부릅니다. Python은 ok입니다.

diff() — 두 문서의 차이

edited = plain.insert_para(section=0, anchor=1, text="새 문단").document
d = plain.diff(edited)
print(d["identical"])                      # False
print(d["package"]["changed"])             # ['Contents/section0.xml']
print(len(d["semantic"]["paragraphs"]), len(d["semantic"]["structure"]))

semantic은 디코드된 모델의 차이(field_values·cells·paragraphs·structure·raw), package는 ZIP 엔트리를 바이트로 비교한 결과(added·removed·changed)입니다. note가 이 두 층의 의미를 한 줄로 설명합니다. 조판 캐시처럼 디코드 모델에 담기지 않는 XML의 차이는 raw로 잡히며 raw_dropped가 생략된 개수를 말합니다.

stamp_plan() — 스탬프할 자리 찾기

plan = hwpforge.Document.open("template.hwpx").stamp_plan()
for c in plan["text"]:
    print(c["path"], c["span"], repr(c["marker"]), c["pattern"], c["guard"])
for c in plan["cells"]:
    print(c["table"], c["at"], [label["normalized"] for label in c["labels"]], c.get("suggested_name"))
print(plan["schema_version"], plan["source_sha256"][:12])

text는 ( )·□ 같은 인라인 표시 후보, cells는 라벨 옆 빈 셀 후보입니다. guard가 있는 후보는 문맥상 채우면 안 될 가능성이 있는 자리(예: 안내문 안의 괄호)라 stamp가 승인 없이는 건드리지 않습니다. source_sha256은 이 계획이 어느 문서에서 나왔는지의 지문이고, 버전 있는 요청(StampRequestV2)에 그대로 들어갑니다. 계획을 실제 스탬프로 바꾸는 규칙은 편집에 있습니다.

to_json() · export_section() — 구조를 읽는 가장 넓은 창

to_json()은 문서 전체를, export_section(section=i)은 한 구역을 JSON으로 내보냅니다. 둘 다 TextResult라 .text(JSON 문자열)와 .report(같은 내용을 dict로 담은 document/section 키 + warnings)를 함께 줍니다. 두 JSON은 서로 다른 스키마입니다 — 전체 문서는 from_json이, 한 구역은 patch가 읽습니다. 자세한 것은 변환과 내보내기와 편집에 있습니다.

편집

모든 편집 메서드는 DocumentResult를 돌려줍니다 — .document가 새 문서, .report가 보고서입니다. 받은 문서는 바뀌지 않으므로 결과를 다시 받아야 합니다.

import hwpforge

doc = hwpforge.Document.open("form.hwpx")
result = doc.fill({"user_email": "kim@example.com"})
doc = result.document          # 이 줄이 없으면 아무것도 바뀌지 않은 것입니다

편집기는 두 부류입니다. 보존 우선(preserve-first) 편집기(fill·patch)는 건드리는 XML 조각만 바꾸고 나머지 바이트를 그대로 두며, 재인코드 편집기(set_cell·stamp·restyle)는 문서를 모델로 디코드해 고친 뒤 패키지를 다시 씁니다. insert_para/delete_para는 구역 XML에서 해당 문단만 넣고 빼지만, 재인코드 편집기와 같은 사전 검사(원본 패키지의 엔트리가 재인코드 뒤에도 남는지, 디코드–인코드 왕복이 안전한지)를 통과해야 합니다. 이 차이가 아래 “한컴 저장 문서” 표를 만듭니다.

한컴 저장 문서에서 되는 것과 안 되는 것

한컴 오피스가 저장한 .hwpx에는 HwpForge 인코더가 재현하지 않는 패키지 엔트리(Preview/PrvText.txt·Preview/PrvImage.png·META-INF/container.rdf)가 들어 있습니다. 네 연산은 이런 문서를 조용히 손상시키는 대신 fail-closed로 거부합니다.

연산한컴 저장 문서결과
fill동작누름틀만 쓰고 나머지 엔트리는 바이트 그대로 보존
patch동작기존 문단·셀의 텍스트만 바꾸고 패키지를 보존
set_cell·insert_para·delete_para·stamp거부INPUT_ENTRIES_NOT_CARRIED 또는 INPUT_NOT_ROUNDTRIP_SAFE
restyle동작재인코드하므로 위 엔트리와 조판 캐시가 사라지며 warnings로 알림
from_json(base=...)부분 승계이미지만 승계하고 위 엔트리는 버려집니다(별도 경고 없음)
hancom = hwpforge.Document.open("hancom.hwpx")
try:
    hancom.set_cell(table=0, at="0,0", text="값")
except hwpforge.HwpForgeError as exc:
    print(exc.code)     # INPUT_ENTRIES_NOT_CARRIED
    print(exc.hint)     # 되는 길: 텍스트는 export_section → 편집 → patch, 누름틀은 fill

따라서 실제 정부 서식을 다룰 때 통하는 길은 둘입니다: 누름틀은 fill, 그 밖의 텍스트는 export_section → 편집 → patch. 구조를 정말 바꿔야 할 때만 from_json(base=...)을 쓰고, 결과를 한컴에서 열어 확인하세요. 이 제한을 푸는 작업이 진행 중이며(GitHub #143), 풀리면 이 표가 바뀝니다.

fill — 누름틀 채우기

form = hwpforge.Document.open("form.hwpx")
result = form.fill({"user_email": "kim@example.com"})
print(result.report["filled"])
# [{'name': 'user_email', 'section': 0, 'previous': '회사 이메일을 입력하세요'}]
print(result.document.read(field="user_email")["fields"][0]["current"])
# kim@example.com
  • 이름이 없는 필드는 FIELD_NOT_FOUND(메시지에 사용 가능한 이름 목록이 붙습니다), 채울 수 없는 필드는 FIELD_NOT_FILLABLE, 빈 문자열은 EMPTY_FIELD_VALUE, 값 맵이 비어 있으면 NO_VALUES입니다. 값은 문자열이어야 하며 다른 형은 TypeError입니다.
  • 같은 이름의 필드가 문서에 둘 이상이면 어느 쪽인지 모호하므로 자동으로 전부 채우지 않고 FIELD_NAME_AMBIGUOUS로 거부합니다(아무것도 쓰지 않습니다). 문서에서 이름을 유일하게 만든 뒤 다시 채우세요.
  • fill은 바꾼 문단의 조판 캐시만 무효화하고 나머지는 보존하므로, 한컴 저장 문서에서도 안전합니다.
try:
    form.fill({"nope": "x"})
except hwpforge.HwpForgeError as exc:
    print(exc.code, "|", exc.message)
    # FIELD_NOT_FOUND | field 'nope' not found; available: [user_email]

patch — 구역 JSON을 편집해 되돌려 넣기

텍스트를 바꾸는 보존 우선 경로입니다. export_section이 낸 JSON의 텍스트를 고쳐 patch로 넣으면 그 구역 XML의 텍스트 슬롯만 치환됩니다.

import json

exported = hancom.export_section(section=0)
section = json.loads(exported.text)          # 키: section_index, section, styles, preservation
patched = hancom.patch(section=0, patch=json.dumps(section, ensure_ascii=False))
print(patched.report)                        # {'section': 0, 'warnings': []}
print(hancom.diff(patched.document)["identical"])   # 바꾼 게 없으면 True
  • JSON의 텍스트만 바꿀 수 있습니다. 문단을 더하거나 빼거나 표 구조·서식 참조를 바꾸면 PATCH_FAILED(구조 변경 감지)로 거부됩니다. 구조 변경은 insert_para/delete_para/set_cell 또는 from_json의 몫입니다.
  • preservation 키는 텍스트 슬롯과 원본 바이트 스팬의 대응표입니다. 손대지 말고 그대로 돌려보내세요.
  • JSON이 구역 스키마에 맞지 않으면 JSON_PARSE_FAILED, 격자 주소가 더 이상 맞지 않으면 PATCH_FAILED입니다. 스키마는 hwpforge.schema(kind="exported-section")으로 받을 수 있습니다.

텍스트를 실제로 바꾸는 예:

section = json.loads(hancom.export_section(section=0).text)
first_run = section["section"]["paragraphs"][0]["runs"][0]
if "text" in first_run:
    first_run["text"] = "바뀐 첫 문단"
doc2 = hancom.patch(section=0, patch=json.dumps(section, ensure_ascii=False)).document

문단의 runs[i]는 텍스트 런이면 text 키를, 표·그림 같은 컨트롤이면 다른 키를 가집니다. 표 셀의 텍스트도 같은 JSON 안에 있으므로(runs[i]["table"]["rows"][r]["cells"][c]["paragraphs"]…) 한컴 저장 문서의 셀 텍스트는 이 길로 바꿉니다.

set_cell — 표 셀에 쓰기

격자 주소(at)나 라벨 기준(right_of·below)으로 셀 하나를 지정하거나, specs로 여럿을 한 번에 씁니다. 두 형태는 배타적입니다.

grid = hwpforge.Document.open("grid.hwpx")
print([(c["row"], c["col"], c["text"]) for c in grid.read(table=0)["table"]["cells"]])
# [(0, 0, '성명'), (0, 1, ''), (1, 0, '비고'), (1, 1, ''), (2, 1, '')]

one = grid.set_cell(table=0, at="0,1", text="홍길동")
print(one.report["results"][0])
# {'table': 0, 'requested': {'row': 0, 'col': 1}, 'anchor': {'row': 0, 'col': 1}, 'resolution': 'exact', 'cleared': False}

by_label = grid.set_cell(table=0, right_of="성명", text="홍길동")

many = grid.set_cell(specs=[
    {"table": 0, "at": {"row": 0, "col": 1}, "text": "홍길동"},
    {"table": 0, "right_of": "비고", "text": "없음"},
])
print([r["resolution"] for r in many.report["results"]])
  • at은 메서드 인자로는 "row,col" 문자열, specs 안에서는 {"row": r, "col": c} 객체입니다(0부터).
  • right_of/below는 그 텍스트를 가진 셀의 오른쪽/아래 셀입니다. 라벨이 없거나 그 방향에 셀이 없으면 CELL_NOT_FOUND, 라벨이 여럿이면 모호함으로 거부됩니다.
  • 병합된 영역을 지정하면 기준 셀로 해석되며 resolution이 그 사실을 말합니다(exact가 아닌 값). 표가 격자 주소를 지원하지 않으면(outline의 addressable 거짓) 거부됩니다.
  • 표 번호는 outline()["outline"]["tables"][i]["ordinal"]이며, 없으면 TABLE_NOT_FOUND.
  • 재인코드 편집기이므로 한컴 저장 문서는 INPUT_ENTRIES_NOT_CARRIED/INPUT_NOT_ROUNDTRIP_SAFE로 거부됩니다. 그 문서의 셀 텍스트는 patch로 바꾸세요.

insert_para / delete_para — 문단 넣고 빼기

plain = hwpforge.Document.open("plain.hwpx")
print([p["text"] for p in plain.read(section=0, paras="0..3")["paragraphs"]["paragraphs"]])
# ['첫째 문단입니다.', '둘째 문단입니다.', '셋째 문단입니다.', '넷째 문단입니다.']

ins = plain.insert_para(section=0, anchor=1, text=["새 문단 하나", "새 문단 둘"])
print(ins.report)                    # {'inserted': 2, 'deleted': 0, 'warnings': []}

above = plain.insert_para(section=0, anchor=1, text="위에", before=True)

dl = ins.document.delete_para(section=0, indexes=[2, 3])
print(dl.report)                     # {'inserted': 0, 'deleted': 2, 'warnings': []}
  • text는 문자열 하나 = 문단 하나, 시퀀스 = 원소마다 문단 하나입니다. 문자열을 글자 단위 시퀀스로 풀지 않습니다. 시퀀스의 원소는 str만 받습니다.
  • anchor/indexes는 그 구역의 최상위 문단 인덱스입니다(표 셀 안의 문단은 세지 않습니다). 범위를 넘으면 PARAGRAPH_OUT_OF_RANGE.
  • 삭제는 전부-아니면-무(all-or-nothing)이며 fail-closed 정책이 있습니다: 구역 속성을 가진 첫 문단(SECTION_PROPERTIES_PARAGRAPH), 책갈피·상호참조·각주 등 참조를 가진 문단, 쪽/단 나눔을 가진 문단(HARD_BREAK_LOSS), 구역을 비우는 삭제는 거부됩니다. 같은 이유로 첫 문단 앞에는 넣을 수 없습니다(INSERT_BEFORE_SECTION_PROPERTIES).
  • 두 연산은 대상 구역 XML만 바꾸고 다른 엔트리는 바이트 그대로 둡니다. 다만 현재는 재인코드 편집기와 같은 사전 검사를 거치므로 한컴 저장 문서를 INPUT_ENTRIES_NOT_CARRIED로 거부합니다(위 표).
try:
    plain.delete_para(section=0, indexes=[0])
except hwpforge.HwpForgeError as exc:
    print(exc.code)      # SECTION_PROPERTIES_PARAGRAPH

stamp — 템플릿에 누름틀 심기

stamp_plan이 찾은 후보를 누름틀로 바꾸는 연산입니다. 흐름은 계획 → 명세(spec) 작성 → 스탬프 이고, 가드 없는 후보는 전부 이름을 붙이거나 "ignore"로 명시해야 합니다 — 하나라도 빠지면 STAMP_CANDIDATE_UNCOVERED로 아무것도 쓰지 않습니다.

template = hwpforge.Document.open("template.hwpx")
plan = template.stamp_plan()

specs = []
for i, candidate in enumerate(plan["text"]):
    spec = dict(candidate)                     # section, path, span, marker 를 그대로 복사
    if i == 0:
        spec["action"] = {"field": {"name": "applicant", "hint": "신청인 이름"}}
    else:
        spec["action"] = "ignore"
    specs.append(spec)

stamped = template.stamp(specs)
print(stamped.report["stamped"])
# [{'name': 'applicant', 'section': 0, 'path': 'paragraphs[0].runs[0].text', 'span': {...}, 'marker': '(   )', 'pattern': 'paren_blank'}]
print(stamped.report["ignored"], stamped.report["skipped_guarded"])
print([(f["name"], f["hint"]) for f in stamped.document.fields()["fields"]])
# [('applicant', '신청인 이름')]

명세 작성 규칙:

  • 텍스트 후보: 계획의 후보 객체를 복사하고 action만 더합니다. section·path·span·marker가 빠지면 거부되고, 여분 키(pattern·guard)는 무시됩니다.
  • 셀 후보: 복사하지 말고 table·at·action으로 새로 만듭니다(label을 붙이면 text는 계획의 labels[].normalized여야 합니다). 후보 객체를 통째로 넘기면 알 수 없는 키로 거부됩니다.
  • action은 "ignore" 또는 {"field": {"name": ..., "hint": ...}}입니다.
  • 버전 있는 요청은 {"schema_version": plan["schema_version"], "source_sha256": plan["source_sha256"], "text": [...], "cells": [...]}이며, 지문이 문서와 다르면 거부됩니다 — 계획을 만든 문서에만 적용된다는 뜻입니다.
request = {
    "schema_version": plan["schema_version"],
    "source_sha256": plan["source_sha256"],
    "text": specs,
}
stamped = template.stamp(request, manifest=False)      # 보고서에서 manifest 생략

보고서의 manifest는 무엇을 어디에 심었는지의 기록(schema_version·source_sha256·output_sha256·fields)이고, stamped/stamped_cells는 적용 단계의 결과입니다. 재인코드 편집기이므로 의미가 손상될 인코딩은 ENCODE_SEMANTIC_LOSS로 거부하며(exc.details에 원인 경고), 한컴 저장 문서는 편집 전 사전 검사에서 거부됩니다.

restyle — 다른 프리셋으로

print([p["name"] for p in hwpforge.templates()["presets"]])   # ['default', 'modern', 'classic', 'latest']
restyled = hancom.restyle(preset="modern")
print(restyled.report["preset"], restyled.report["paragraphs"])
print([w["code"] for w in restyled.report["warnings"]])

문서를 다시 인코드하므로 조판 캐시와 인코더가 모르는 엔트리는 사라지고 그 사실이 warnings로 옵니다. 없는 프리셋은 PRESET_NOT_FOUND, 의미 손상은 ENCODE_SEMANTIC_LOSS입니다.

편집 뒤에 할 일

  1. result.report["warnings"]를 읽습니다. 비어 있어야 정상이고, LAYOUT_CACHE_DROPPED 같은 코드는 한컴에서 다시 저장해야 쪽 배치가 복원된다는 뜻입니다.
  2. result.document.validate()["ok"]로 모델이 유효한지 봅니다.
  3. original.diff(result.document)로 바뀐 범위가 의도와 같은지 봅니다 — package["changed"]가 대상 구역 하나뿐인지가 좋은 검사입니다.
  4. save 합니다. 파일명은 .hwpx로.

변환과 내보내기

convert_md — Markdown에서 문서 만들기

import hwpforge

source = """---
title: 보고서
---

# 개요

본문입니다.

| 항목 | 값 |
| -- | -- |
| 예산 | 1,000 |
"""
result = hwpforge.convert_md(source, preset="classic")
print(result.report["sections"], result.report["paragraphs"])
print(result.document.inspect()["metadata"]["title"])     # 보고서
result.document.save("report.hwpx")
  • Markdown은 GFM이고, 맨 앞의 YAML frontmatter가 문서 메타데이터와 스타일을 정합니다(title 등). 스타일 문법은 Markdown에서 HWPX 로를 보세요.
  • preset은 hwpforge.templates()가 나열하는 이름 중 하나입니다(default·modern·classic·latest). 없는 이름은 PRESET_NOT_FOUND.
  • 이미지 ![…](path)는 base_dir를 기준으로 읽습니다. base_dir를 주지 않으면 파일 이미지는 버려지고 보고서에 남습니다. base_dir 밖의 경로는 읽지 않습니다.
result = hwpforge.convert_md("# 그림\n\n![도표](chart.png)\n")     # base_dir 없음
print(result.report["assets"])
# [{'kind': 'dropped', 'occurrence': {'paragraph': 1, 'run': 0}, 'reason': 'no_base_dir'}]
print([w["code"] for w in result.report["warnings"]])              # ['IMAGE_EMBED_SKIPPED']

assets의 각 항목은 kind로 구분됩니다: embedded(key·format 포함), dropped(reason 포함), remote(URL은 내려받지 않습니다).

to_md — 문서를 Markdown으로

doc = hwpforge.Document.open("hancom.hwpx")
styled = doc.to_md()                       # mode="styled": 스타일 frontmatter 유지
lossy = doc.to_md(mode="lossy")            # Markdown 이 못 담는 것은 버리고 warnings 로 알림
print(lossy.text[:200])
print(lossy.report["mode"], list(lossy.report["images"].keys()))
  • mode="lossless"는 무언가를 잃어야 하면 ENCODE_FAILED로 거부합니다.
  • 보고서의 images는 문서가 참조한 이미지의 {키: 바이트}입니다. Markdown 텍스트가 그 키를 가리키므로, 파일로 저장할 때 같은 이름으로 옆에 써 두면 됩니다.

to_json · from_json — 전체 문서를 JSON으로

import json

exported = doc.to_json()                            # styles=True 가 기본
document = json.loads(exported.text)                # 키: document, styles
print(exported.report["document"] == document)      # 같은 내용을 dict 로도 줍니다

rebuilt = hwpforge.from_json(exported.text, base=doc)
print(rebuilt.report["paragraphs"], [w["code"] for w in rebuilt.report["warnings"]])
# 3 ['LAYOUT_CACHE_DROPPED']
  • to_json의 JSON은 문서 전체와 style store이며, 표 셀에 격자 주소가 붙어 있습니다. styles=False는 구조만 내보냅니다.
  • from_json은 그 JSON을 다시 인코드해 새 문서를 만듭니다. base를 주면 style store(스타일 없이 내보낸 JSON 일 때)와 이미지 바이너리를 그 문서에서 가져옵니다. base 없이 이미지가 있는 문서를 재구성하면 이미지가 빠집니다.
  • 생성은 fail-closed가 아닙니다: 의미가 손실돼도 거부하지 않고 warnings에 실어 돌려줍니다. 위의 LAYOUT_CACHE_DROPPED는 원본이 가진 줄 조판 캐시를 재인코드가 싣지 않는다는 뜻입니다(한컴에서 다시 저장하면 복원).
  • 한컴 저장 문서의 Preview/*·container.rdf 엔트리는 base를 줘도 승계되지 않습니다(경고 없음). 텍스트만 바꿀 거라면 export_section → patch가 문서를 온전히 보존합니다.
  • JSON이 문서가 아니면 JSON_PARSE_FAILED. 스키마는 hwpforge.schema(kind="exported-document").

전체 JSON(document/styles)과 구역 JSON(section_index/section/styles/preservation)은 서로 다른 스키마입니다 — 전자는 from_json, 후자는 patch가 읽습니다.

convert_hwp5 — 옛 .hwp를 HWPX로

with open("old.hwp", "rb") as handle:
    converted = hwpforge.convert_hwp5(handle.read())
print(len(converted.document), [w["code"] for w in converted.report["warnings"]])
converted.document.save("old.hwpx")
  • 인자는 파일 경로가 아니라 바이트입니다(Document.open은 HWPX 전용입니다). 읽을 수 없는 바이트는 HWP5_DECODE_FAILED.
  • 보고서의 warnings에 HWP5가 가졌지만 HWPX로 옮기지 못한 것이 전부 나옵니다.
  • carry_layout_cache=True는 한컴이 계산해 둔 줄 조판 캐시를 함께 옮깁니다. PDF 렌더가 한컴의 쪽 나눔을 그대로 재현하려면 이 캐시가 필요합니다(아래). 캐시를 옮긴 HWPX는 PDF 재생·비교용이며, 한컴에서 다시 열어 편집할 문서로 취급하지 마세요.

to_pdf — PDF 렌더

렌더는 문서에 저장된 조판을 재생합니다(다시 계산하지 않습니다). 그래서 조판 캐시가 있는 문서만 렌더 대상입니다: 한컴이 저장한 HWPX, 그리고 convert_hwp5(..., carry_layout_cache=True)로 캐시를 옮긴 HWPX. Markdown·JSON에서 생성한 문서는 캐시가 없어 PDF_RENDER_FAILED입니다.

캐시가 있어도 아직 재생하지 못하는 내용이 있습니다. 한컴이 저장한 문서라도 아래 경우는 실패하고, 한컴에서 다시 저장해도 풀리지 않습니다.

cause["code"]언제예
INVALID_CACHE문단 안에 글자가 아닌 요소(누름틀·메모·각주·수식·하이퍼링크·상호참조·차트·도형 등)가 있음누름틀이 든 서식 (clickhere_filled.hwpx)
UNSUPPORTED_CONTENT아직 지원하지 않는 배치 — cause["kind"]가 무엇인지 말함kind == "non-default table position" (table_20_real_world_ministry_stress.hwpx)

누름틀 서식처럼 이 경우에 걸리는 문서는 한컴에서 PDF로 저장하세요.

with open("old.hwp", "rb") as handle:
    carried = hwpforge.convert_hwp5(handle.read(), carry_layout_cache=True).document

try:
    pdf = carried.to_pdf(font_dirs=["/Library/Fonts/Hancom"])
    print(pdf.report["pages"])
    with open("old.pdf", "wb") as out:
        out.write(pdf.data)
except hwpforge.HwpForgeError as exc:
    print(exc.code, exc.cause)
    # PDF_RENDER_FAILED {'stage': 'render', 'code': 'FONT_UNRESOLVED'}      ← 글꼴을 못 찾음
    # PDF_RENDER_FAILED {'stage': 'render', 'code': 'NO_RENDERABLE_CACHE', 'location': 's0'}  ← 캐시 없음
    # PDF_RENDER_FAILED {'stage': 'render', 'code': 'INVALID_CACHE'}        ← 재생할 수 없는 요소가 든 문단
  • 글꼴은 기본 fail-closed입니다: 문서가 이름 붙인 글꼴 face를 font_dirs 안에서 찾지 못하면 추측하지 않고 실패합니다(cause["code"] == "FONT_UNRESOLVED"). degraded=True는 대체 글꼴로 렌더하며, 결과의 모양이 달라집니다.
  • discovery는 font_dirs 밖을 더 볼지의 선택입니다: "explicit"(기본, 결정적) · "hancom"(한컴 설치 글꼴 위치) · "platform"(OS 글꼴).
  • font_dirs는 시퀀스여야 합니다. 문자열 하나를 주면 한 글자짜리 디렉터리들로 읽히는 대신 TypeError로 거부됩니다.
  • partial_cache_reject=True는 캐시가 일부만 있는 문서를 다시 배치하지 않고 거부합니다.
  • 실패의 두 번째 분류는 exc.cause에 옵니다(stage·code·kind·location). exc.hint는 원인과 상관없이 같은 문장이므로, 무엇을 할지는 cause["code"]로 가르세요: FONT_UNRESOLVED는 font_dirs·discovery, NO_RENDERABLE_CACHE·MISSING_LAYOUT_CACHE는 한컴에서 다시 저장, INVALID_CACHE·UNSUPPORTED_CONTENT는 위 표.

templates · schema

for preset in hwpforge.templates()["presets"]:
    print(preset["name"], preset["description"], preset["font"], preset["page_size"])

schema = hwpforge.schema(kind="exported-section")   # "document" | "exported-document" | "exported-section"
print(schema["title"], list(schema["properties"])[:4])

두 함수의 결과에는 warnings가 없습니다. schema는 JSON Schema dict이며 키는 스키마 자신의 것입니다.

결과·경고·오류

결과 객체 세 가지

값을 만들어 내는 연산은 값과 보고서를 한 객체로 돌려줍니다. 셋 다 frozen dataclass입니다.

타입필드돌려주는 연산
DocumentResult[R]document: Document, report: Rfill·set_cell·patch·insert_para·delete_para·stamp·restyle, convert_md·from_json·convert_hwp5
TextResult[R]text: str, report: Rto_json·export_section·to_md
BytesResult[R]data: bytes, report: Rto_pdf

검사 연산(inspect·outline·fields·validate·read·diff·stamp_plan)과 templates·schema는 보고서 dict만 돌려줍니다.

import hwpforge

result = hwpforge.convert_md("# 제목\n\n본문")
result.document          # Document
result.report            # {'sections': 1, 'paragraphs': ..., 'assets': [], 'warnings': []}
document, report = result.document, result.report      # 필드 이름으로 풀어 쓰는 편이 안전합니다

보고서와 타입 힌트

보고서는 TypedDict입니다. 키 이름은 연산 계층의 보고서 구조체 필드명과 같고, hwpforge._hwpforge 스텁(_hwpforge.pyi)에 전부 선언돼 있습니다. 편집기에서 자동 완성과 타입 검사를 받으려면 그 이름을 가져오세요(스텁 모듈 자체는 비공개이지만 타입 이름은 안정적인 편입니다).

from typing import TYPE_CHECKING

if TYPE_CHECKING:
    from hwpforge._hwpforge import FillReport, InspectReport

def summarize(report: "InspectReport") -> str:
    return f"{report['sections']} sections, {report['paragraphs']} paragraphs"

NotRequired로 선언된 키(예: InspectReport["styles"], StampReport["manifest"], 경고의 hint)는 없을 수 있으니 .get()으로 읽으세요. 그 밖의 키는 항상 있고, 값이 없으면 None입니다.

warnings — 조용한 손실을 드러내는 채널

templates()·schema()를 뺀 모든 보고서에 warnings: list[{code, message, hint?}]가 있습니다. 실패가 아니라 “성공했지만 이것은 보존되지 않았다“는 신호이므로, 저장하기 전에 읽는 습관이 필요합니다.

doc = hwpforge.Document.open("hancom.hwpx")
rebuilt = hwpforge.from_json(doc.to_json().text, base=doc)
for warning in rebuilt.report["warnings"]:
    print(warning["code"], "-", warning["message"])
    if warning.get("hint"):
        print("  →", warning["hint"])
# LAYOUT_CACHE_DROPPED - ...

자주 보게 되는 코드:

코드뜻어디서
LAYOUT_CACHE_DROPPED원본의 줄 조판 캐시를 재인코드가 싣지 않았다 — 한컴에서 다시 저장하면 복원from_json·restyle 등 재인코드 경로
IMAGE_EMBED_SKIPPEDMarkdown의 이미지를 읽지 못해 뺐다(base_dir 없음, 경로 밖, 미지원 형식)convert_md
디코더 경고문서가 가졌지만 모델이 담지 못한 것(예: 알 수 없는 컨트롤)디코드하는 모든 연산

같은 경고를 validate()의 warnings로도 볼 수 있으므로, 문서를 처음 받았을 때 한 번 validate()를 부르는 것이 좋은 시작입니다.

HwpForgeError — 하나의 예외, 다섯 속성

입력을 받아들인 뒤 연산이 거부하면 HwpForgeError를 던집니다. str(exc)는 "CODE: message"(힌트가 있으면 둘째 줄에 hint: …)입니다.

속성타입내용
codestr안정된 실패 코드. 분기는 이 문자열로
messagestr무엇이 잘못됐는지 한 문장
hintstr | None어떻게 해야 하는지 — 연산이 알 때만
causedict | None두 번째 분류. 현재는 PDF_RENDER_FAILED만 {"stage", "code", "kind"?, "location"?}를 가짐
detailsdict | None구조화된 페이로드. ENCODE_SEMANTIC_LOSS는 {"warnings": [...], "others": [...]}
try:
    doc.set_cell(table=0, at="9,9", text="값")
except hwpforge.HwpForgeError as exc:
    if exc.code == "CELL_NOT_FOUND":
        ...
    elif exc.code in ("INPUT_ENTRIES_NOT_CARRIED", "INPUT_NOT_ROUNDTRIP_SAFE"):
        ...     # 한컴 저장 문서 — patch 경로로
    else:
        raise

code는 CLI·MCP와 대체로 같지만 같은 실패라도 창구마다 동결된 문자열이 다를 수 있습니다(예: CLI DECODE_FAILED와 MCP DECODE_ERROR). Python 코드는 이 패키지의 계약이므로 이 문서의 표를 기준으로 삼으세요.

다른 예외

예외언제
TypeError / ValueError인자 변환 실패 — fill 값에 정수, font_dirs에 문자열 하나, Document(…)에 bytes 아닌 것
OSErrorDocument.open·save의 파일 I/O
AttributeErrorDocument 속성 대입·삭제(불변)
pyo3_runtime.PanicException내부 Rust panic — 버그이니 이슈로 알려 주세요

코드 표

각 연산이 자기 특성 실패로 내는 코드입니다. HWPX 문서를 디코드하는 연산은 디코드 실패를 DECODE_FAILED로 냅니다. 다만 fill·set_cell·insert_para·delete_para·stamp·stamp_plan은 디코드 실패도 자기 codec 코드(FILL_FAILED·SET_CELL_CODEC_FAILED·STRUCTURAL_CODEC·STAMP_CODEC_FAILED)로 내므로 DECODE_FAILED는 나오지 않습니다. 코드는 hwpforge::foundation::diagnostics::OpsCode의 문자열 그대로이며, 아래 표는 Python이 부르는 연산에서 실제로 나올 수 있는 것만 적었습니다.

연산코드
Document.openINPUT_TOO_LARGE
그 밖에 HWPX를 디코드하는 연산DECODE_FAILED
readREAD_TARGET_REQUIRED · READ_PARAS_INVALID · READ_PARAS_WITHOUT_SECTION · READ_SECTION_OUT_OF_RANGE · READ_PARA_RANGE_INVALID · READ_TABLE_OUT_OF_RANGE · READ_FIELD_NOT_FOUND · TABLE_GRID_INVALID
export_sectionSECTION_OUT_OF_RANGE · GRID_ADDR_PROJECTION_FAILED · JSON_SERIALIZE_FAILED
to_jsonGRID_ADDR_PROJECTION_FAILED · JSON_SERIALIZE_FAILED
to_mdINVALID_INPUT (알 수 없는 mode) · ENCODE_FAILED (mode="lossless") · VALIDATION_FAILED
to_pdfINVALID_DISCOVERY · UNRECOGNIZED_FORMAT · HWP5_DECODE_FAILED · HWP5_CONVERT_FAILED · VALIDATION_FAILED · PDF_RENDER_FAILED (+ cause)
fillNO_VALUES · FIELD_NOT_FOUND · FIELD_NAME_AMBIGUOUS · FIELD_NOT_FILLABLE · EMPTY_FIELD_VALUE · FILL_FAILED
set_cellINVALID_SET_CELL_ARGS · INVALID_SET_CELL_MAP · TABLE_NOT_FOUND · TABLE_GRID_INVALID · CELL_NOT_FOUND · CELL_LABEL_AMBIGUOUS · CELL_HAS_NON_TEXT_CONTENT · CELL_TARGET_DUPLICATE · CELL_TARGET_CONFLICT · SET_CELL_CODEC_FAILED · INPUT_ENTRIES_NOT_CARRIED · INPUT_NOT_ROUNDTRIP_SAFE · ENCODE_SEMANTIC_LOSS
patchJSON_PARSE_FAILED · GRID_ADDR_INVALID · SECTION_OUT_OF_RANGE · SECTION_INDEX_MISMATCH · PATCH_FAILED
insert_paraINSERT_TEXT_REQUIRED · MULTI_PARAGRAPH_TEXT · SECTION_OUT_OF_RANGE · PARAGRAPH_OUT_OF_RANGE · INSERT_BEFORE_SECTION_PROPERTIES · STRUCTURAL_CODEC · INPUT_ENTRIES_NOT_CARRIED · INPUT_NOT_ROUNDTRIP_SAFE · SELF_VERIFY_FAILED · SPAN_COUNT_MISMATCH
delete_paraDELETE_NO_TARGET · SECTION_OUT_OF_RANGE · PARAGRAPH_OUT_OF_RANGE · DUPLICATE_TARGET · REFERENCE_STRANDED · HARD_BREAK_LOSS · EMPTY_SECTION · SECTION_PROPERTIES_PARAGRAPH · STRUCTURAL_CODEC · INPUT_ENTRIES_NOT_CARRIED · INPUT_NOT_ROUNDTRIP_SAFE · SELF_VERIFY_FAILED · SPAN_COUNT_MISMATCH
stampINVALID_STAMP_MAP · STAMP_SPEC_STALE · STAMP_MARKER_MISMATCH · STAMP_SPEC_DUPLICATE · STAMP_NAME_EMPTY · STAMP_NAME_DUPLICATE · STAMP_NAME_COLLISION · STAMP_CANDIDATE_UNCOVERED · STAMP_CELL_NOT_ANCHOR · STAMP_CELL_NOT_EMPTY · STAMP_LABEL_DRIFT · STAMP_CELL_NOT_CANDIDATE · STAMP_CELL_TARGET_DUPLICATE · STAMP_SOURCE_HASH_MISMATCH · STAMP_MANIFEST_INVARIANT · STAMP_DELTA_MISMATCH · STAMP_CODEC_FAILED · TABLE_NOT_FOUND · TABLE_GRID_INVALID · INPUT_ENTRIES_NOT_CARRIED · INPUT_NOT_ROUNDTRIP_SAFE · ENCODE_SEMANTIC_LOSS
stamp_planSTAMP_CODEC_FAILED
restylePRESET_NOT_FOUND · NO_FONTS · VALIDATION_FAILED · ENCODE_FAILED · ENCODE_SEMANTIC_LOSS
convert_mdPRESET_NOT_FOUND · MD_DECODE_FAILED · STYLE_STORE_FAILED · STYLE_REBIND_FAILED · VALIDATION_FAILED · ENCODE_FAILED
from_jsonJSON_PARSE_FAILED · GRID_ADDR_INVALID · VALIDATION_FAILED · ENCODE_FAILED
convert_hwp5HWP5_DECODE_FAILED · HWP5_CONVERT_FAILED
schemaINVALID_INPUT (알 수 없는 kind)

templates·schema를 뺀 대부분의 연산에서 UPSTREAM_UNMAPPED(이 버전이 아직 분류하지 못한 상류 오류)와 INTERNAL_INVARIANT(라이브러리 불변식 위반 — 버그)가 나올 수 있습니다. 둘 다 이슈로 알려 주세요.

fail-closed 거부 읽기 — ENCODE_SEMANTIC_LOSS

재인코드 편집기(stamp·restyle)는 인코딩이 의미를 잃을 것 같으면 바이트를 아예 만들지 않고 거부합니다. 무엇이 문제였는지는 details에 있습니다.

try:
    doc.restyle(preset="modern")
except hwpforge.HwpForgeError as exc:
    if exc.code == "ENCODE_SEMANTIC_LOSS":
        for w in exc.details["warnings"]:      # 거부를 일으킨 의미 손상
            print(w["code"], w["message"])
        for w in exc.details["others"]:        # 같은 인코딩이 낸 그 밖의 경고
            print("also:", w["code"])

예외를 직렬화·복제할 때

HwpForgeError는 copy·pickle이 다섯 속성을 그대로 보존하도록 __reduce__를 정의합니다. 멀티프로세스 워커에서 예외를 부모로 넘겨도 code·details가 살아 있습니다.

레시피

양식 채우기 파이프라인

정부 서식처럼 한컴이 저장한 양식을 값으로 채워 저장하는 가장 흔한 흐름입니다. 누름틀은 fill, 나머지는 검증·저장.

import hwpforge

def fill_form(src: str, dst: str, values: dict[str, str]) -> list[str]:
    doc = hwpforge.Document.open(src)

    available = {f["name"] for f in doc.fields()["fields"] if f["fillable"]}
    unknown = set(values) - available
    if unknown:
        raise ValueError(f"양식에 없는 필드: {sorted(unknown)}")

    result = doc.fill(values)
    problems = [f"{w['code']}: {w['message']}" for w in result.report["warnings"]]

    check = result.document.validate()
    if not check["ok"]:
        problems += [f"{e['code']}: {e['message']}" for e in check["errors"]]

    result.document.save(dst)
    return problems

print(fill_form("form.hwpx", "form-filled.hwpx", {"user_email": "kim@example.com"}))
  • 필드 이름을 미리 대조하면 FIELD_NOT_FOUND를 예외가 아니라 목록으로 다룰 수 있습니다.
  • 값이 비어 있으면 EMPTY_FIELD_VALUE입니다. 지우고 싶은 필드는 값 목록에서 빼세요.
  • fill은 한컴 저장 문서를 보존하므로 결과를 한컴에서 열면 그대로 보입니다. 재인코드가 없어 조판 캐시도 건드린 문단만 무효화됩니다.

표를 읽어 CSV로

import csv
import hwpforge

def tables_to_csv(path: str, out_prefix: str) -> int:
    doc = hwpforge.Document.open(path)
    tables = doc.outline()["outline"]["tables"]
    for t in tables:
        view = doc.read(table=t["ordinal"])["table"]
        grid = [[""] * view["cols"] for _ in range(view["rows"])]
        for cell in view["cells"]:
            grid[cell["row"]][cell["col"]] = cell["text"]     # 병합 영역은 기준 셀에만 텍스트가 있습니다
        with open(f"{out_prefix}-{t['ordinal']}.csv", "w", newline="", encoding="utf-8") as f:
            csv.writer(f).writerows(grid)
    return len(tables)

print(tables_to_csv("grid.hwpx", "grid"))

텍스트만 바꾸기 — 한컴 저장 문서에서도 되는 길

import json
import hwpforge

def replace_text(doc: hwpforge.Document, section: int, old: str, new: str) -> hwpforge.Document:
    exported = doc.export_section(section=section)
    payload = json.loads(exported.text)

    def walk(node):
        if isinstance(node, dict):
            if isinstance(node.get("text"), str) and old in node["text"]:
                node["text"] = node["text"].replace(old, new)
            for value in node.values():
                walk(value)
        elif isinstance(node, list):
            for item in node:
                walk(item)

    walk(payload["section"])
    return doc.patch(section=section, patch=json.dumps(payload, ensure_ascii=False)).document

doc = hwpforge.Document.open("hancom.hwpx")
changed = replace_text(doc, 0, "표", "테이블")
print(doc.diff(changed)["package"]["changed"])

patch는 텍스트 슬롯만 치환하므로 문단 수·표 구조·서식 참조는 건드리면 안 됩니다. 위 함수처럼 text 값만 바꾸면 PATCH_FAILED가 나지 않습니다.

디렉터리 배치 처리

from pathlib import Path
import hwpforge

def inventory(root: str) -> list[dict]:
    rows = []
    for path in sorted(Path(root).glob("*.hwpx")):
        doc = hwpforge.Document.open(path)
        try:
            report = doc.inspect()
        except hwpforge.HwpForgeError as exc:
            rows.append({"file": path.name, "error": exc.code})
            continue
        rows.append({
            "file": path.name,
            "title": report["metadata"]["title"],
            "paragraphs": report["paragraphs"],
            "tables": report["tables"],
            "fields": len(report["fields"]),
            "warnings": len(report["warnings"]),
        })
    return rows

for row in inventory("."):
    print(row)

문서는 open 시점에 디코드되지 않으므로 파일마다 inspect에서 실패를 잡으면 됩니다. CPU를 더 쓰고 싶으면 concurrent.futures.ProcessPoolExecutor로 파일 단위로 나누세요 — HwpForgeError는 pickle이 되므로 워커의 실패가 code 째로 부모에 돌아옵니다.

HWP를 PDF로

옛 .hwp를 한컴이 계산한 쪽 나눔 그대로 PDF로 만드는 길입니다. 글꼴 디렉터리는 문서가 이름 붙인 face가 실제로 있는 곳이어야 합니다.

import hwpforge

def hwp_to_pdf(src: str, dst: str, font_dirs: list[str]) -> int:
    with open(src, "rb") as handle:
        converted = hwpforge.convert_hwp5(handle.read(), carry_layout_cache=True)
    pdf = converted.document.to_pdf(font_dirs=font_dirs)
    with open(dst, "wb") as out:
        out.write(pdf.data)
    return pdf.report["pages"]

try:
    print(hwp_to_pdf("old.hwp", "old.pdf", ["/Library/Fonts/Hancom"]))
except hwpforge.HwpForgeError as exc:
    print(exc.code, exc.cause and exc.cause["code"])     # FONT_UNRESOLVED 면 font_dirs 를 확인

글꼴을 구할 수 없고 모양이 달라져도 괜찮다면 to_pdf(font_dirs=[...], degraded=True)로 대체 글꼴 렌더를 받을 수 있습니다. carry_layout_cache=True로 변환한 HWPX는 PDF 용이며 한컴에서 다시 열어 편집할 문서로 쓰지 마세요.

Markdown으로 문서 생성 후 검토

import hwpforge

result = hwpforge.convert_md(open("report.md", encoding="utf-8").read(), preset="modern", base_dir=".")
for asset in result.report["assets"]:
    if asset["kind"] == "dropped":
        print("이미지 누락:", asset["occurrence"], asset["reason"])
result.document.save("report.hwpx")

back = result.document.to_md(mode="lossy")
print(back.text[:300])

생성한 문서는 조판 캐시가 없으므로 to_pdf는 되지 않습니다. 한컴에서 열어 저장하면 캐시가 생기고, 재생할 수 없는 요소가 없다면 그때부터 렌더됩니다(to_pdf 표).

pip 없는 호스트에 배포하기

NAS나 잠긴 서버처럼 pip이 없는 곳에서는 wheel을 풀어서 씁니다. 한 번만 하면 되는 절차입니다.

  1. 인터넷이 되는 곳에서 PyPI files 페이지에서 대상 플랫폼의 wheel을 받습니다 — Linux x86_64 면 hwpforge-<version>-cp39-abi3-manylinux_2_28_x86_64.whl. sha256도 그 페이지에 있습니다.
  2. 대상 호스트의 조건을 확인합니다: CPython 3.9 이상, glibc 2.28 이상(ldd --version).
  3. 파일을 옮겨 풉니다.
python3 -c "import zipfile; zipfile.ZipFile('hwpforge-<version>-cp39-abi3-manylinux_2_28_x86_64.whl').extractall('/opt/hf')"
PYTHONPATH=/opt/hf python3 -c "import hwpforge; print(hwpforge.__version__)"

런타임 의존성이 없으므로 이것으로 끝입니다. 스크립트에서는 sys.path.insert(0, "/opt/hf")로도 같은 효과를 냅니다.

문서 두 판을 비교해 검토 보고 만들기

import hwpforge

base = hwpforge.Document.open("form.hwpx")
revised = base.fill({"user_email": "kim@example.com"}).document

d = base.diff(revised)
print("동일:", d["identical"])
print("바뀐 엔트리:", d["package"]["changed"])
for change in d["semantic"]["field_values"]:
    print("필드:", change)
for change in d["semantic"]["paragraphs"]:
    print("문단:", change)

semantic의 다섯 목록(field_values·cells·paragraphs·structure·raw)은 각각 누름틀 값·셀 텍스트·문단 텍스트·구조·모델 밖 XML(조판 캐시 등)의 차이입니다. 편집이 의도한 범위만 건드렸는지 보는 가장 빠른 검사는 package["changed"]가 대상 구역 XML 하나뿐인지입니다.

API 요약

패키지의 공개 표면은 hwpforge.__all__의 열한 이름입니다: Document · DocumentResult · TextResult · BytesResult · HwpForgeError · convert_md · from_json · convert_hwp5 · templates · schema · __version__. 그 밖의 모듈(hwpforge._hwpforge 등)은 비공개이며 릴리스 사이에 바뀔 수 있습니다.

Document

생성과 바이트

이름시그니처설명
Document.open(path) -> Document파일에서. 100 MB 상한, 넘으면 INPUT_TOO_LARGE. I/O 실패는 OSError
Document.from_bytes(data: bytes) -> Document메모리의 바이트에서. 상한 없음
Document(data)(data: bytes)생성자. bytes 아니면 TypeError
to_bytes() -> bytesHWPX 패키지 바이트 그대로
save(path) -> None항상 HWPX로 씀. 기존 파일은 교체
bytes(doc) · len(doc) · == · hash · repr바이트 기준 값 의미. repr은 Document(<n> bytes)

검사 (보고서 dict 반환)

메서드시그니처보고서 키
inspect(*, styles: bool = False) -> InspectReportmetadata·sections·paragraphs·tables·images·charts·fields·section_details·styles?·warnings
outline() -> OutlineReportoutline{title, sections, headings, tables, fields, bookmarks}·warnings
fields() -> FieldsReportfields[{name, hint, current, section, fillable}]·warnings
validate() -> ValidateReportok·sections·paragraphs·errors·warnings
read(*, section=None, paras=None, table=None, field=None) -> ReadReportparagraphs·table·fields(하나만 채워짐)·warnings
diff(revised: Document) -> DiffReportidentical·note·semantic{…}·package{added, removed, changed}·warnings
stamp_plan() -> StampPlanReportschema_version·source_sha256·text·cells·skipped_tables·warnings

내보내기

메서드시그니처반환
to_json(*, styles: bool = True) -> TextResult[ToJsonReport].text JSON, .report{document, warnings}
export_section(*, section: int, styles: bool = True) -> TextResult[ExportSectionReport].text JSON, .report{section, warnings}
to_md(*, mode: "styled" | "lossy" | "lossless" = "styled") -> TextResult[ToMdReport].text Markdown, .report{mode, images, warnings}
to_pdf(*, font_dirs=(), discovery="explicit", degraded=False, partial_cache_reject=False) -> BytesResult[ToPdfReport].data PDF, .report{pages, warnings}

편집 (DocumentResult 반환 — .document 새 문서, .report 보고서)

메서드시그니처보고서 키
fill(values: Mapping[str, str])filled[{name, section, previous}]·warnings
set_cell(*, table=None, at=None, right_of=None, below=None, text=None, specs=None)results[{table, requested, anchor, resolution, cleared}]·warnings
patch(*, section: int, patch: str)section·warnings
insert_para(*, section: int, anchor: int, text: str | Sequence[str], before: bool = False)inserted·deleted·warnings
delete_para(*, section: int, indexes: Sequence[int])inserted·deleted·warnings
stamp(request: StampRequest, *, manifest: bool = True)manifest?·stamped·stamped_cells·ignored·skipped_guarded·warnings
restyle(*, preset: str)preset·sections·paragraphs·warnings

모듈 함수

함수시그니처반환
convert_md(text: str, *, preset: str = "default", base_dir=None)DocumentResult[ConvertMdReport] — sections·paragraphs·assets·warnings
from_json(text: str, *, base: Document | None = None)DocumentResult[EncodeReport] — paragraphs·warnings
convert_hwp5(data: bytes, *, carry_layout_cache: bool = False)DocumentResult[ConvertHwp5Report] — warnings
templates()TemplatesReport — presets[{name, description, font, page_size}]
schema(*, kind: "document" | "exported-document" | "exported-section" = "document")JSON Schema dict

결과 객체와 예외

타입필드/속성
DocumentResult[R]document: Document, report: R
TextResult[R]text: str, report: R
BytesResult[R]data: bytes, report: R
HwpForgeErrorcode: str, message: str, hint: str | None, cause: dict | None, details: dict | None

입력 형태

이름모양
CellSpec{"table": int, "text": str} + at: {"row", "col"} | right_of: str | below: str 중 하나
StampSpec{"section", "path", "span": {"start", "end"}, "marker", "action"} — action은 "ignore" 또는 {"field": {"name", "hint"}}
StampRequestSequence[StampSpec] 또는 {"schema_version", "source_sha256", "text"?: [...], "cells"?: [...]}
paras"시작..끝", 양끝 포함
at (메서드 인자)"row,col" 문자열, 0부터

상수

이름값
hwpforge.__version__설치된 패키지 버전
hwpforge._hwpforge.MAX_FILE_SIZEDocument.open의 상한, 104857600 (100 MB)

HWPX 인코딩/디코딩

HWPX 포맷 소개

HWPX는 한글과컴퓨터의 공개 문서 표준(KS X 6101, OWPML)입니다. 내부 구조는 ZIP 컨테이너 안에 XML 파일들이 담긴 형태로, Microsoft DOCX와 유사합니다.

주요 구성 파일:

  • mimetype — 포맷 식별자
  • Contents/header.xml — 스타일 정의 (폰트, 문단 모양, 글자 모양)
  • Contents/section0.xml, section1.xml, … — 본문 내용
  • BinData/ — 이미지 등 바이너리 파일들
  • Chart/ — 차트 XML (OOXML xmlns:c 형식)

hwpforge-smithy-hwpx 크레이트가 이 포맷의 인코드/디코드를 담당합니다.

디코딩: HWPX 파일 읽기

HwpxDecoder::decode_file()로 .hwpx 파일을 HwpxDocument로 읽습니다.

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwpx::HwpxDecoder;

let result = HwpxDecoder::decode_file("document.hwpx").unwrap();

// 섹션 수 확인
println!("섹션 수: {}", result.document.sections().len());

// 첫 번째 섹션의 문단 수
let section = &result.document.sections()[0];
println!("문단 수: {}", section.paragraphs.len());
}

HwpxDocument 결과 구조

HwpxDecoder::decode_file()은 HwpxDocument를 반환합니다. 네 가지 필드로 구성됩니다. HwpxDocument는 #[non_exhaustive]라서 구조 분해(let HwpxDocument { .. } = ...) 대신 필드로 접근합니다.

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwpx::{HwpxDecoder, HwpxDocument};

let result: HwpxDocument = HwpxDecoder::decode_file("document.hwpx").unwrap();

// result.document: Document<Draft> — 문서 DOM (섹션/문단/런 트리)
// result.style_store: HwpxStyleStore — 폰트, 글자 모양, 문단 모양, 스타일
// result.image_store: ImageStore — 임베드된 이미지 바이너리 데이터
// result.warnings: Vec<DecodeWarning> — 디코드 중 표면화된 비치명 경고
for warning in &result.warnings {
    println!("경고: {warning:?}");
}
}
필드타입설명
documentDocument<Draft>섹션, 문단, 런 트리
style_storeHwpxStyleStore폰트/글자모양/문단모양/스타일
image_storeImageStore이미지 바이너리 저장소
warningsVec<DecodeWarning>디코드 중 표면화된 비치명 경고

메타데이터 접근

디코딩된 문서에서 metadata()로 제목, 작성자 등의 메타데이터에 접근합니다.

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwpx::HwpxDecoder;

let result = HwpxDecoder::decode_file("document.hwpx").unwrap();
let meta = result.document.metadata();

if let Some(title) = &meta.title {
    println!("제목: {}", title);
}
if let Some(author) = &meta.author {
    println!("작성자: {}", author);
}
if let Some(created) = &meta.created {
    println!("작성일: {}", created);
}
}

전체 메타데이터 필드 목록과 사용법은 메타데이터 가이드를 참고하세요.

인코딩: Core → HWPX

HwpxEncoder::encode()로 Document<Validated>를 HWPX 바이트 벡터로 직렬화합니다.

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwpx::{HwpxDecoder, HwpxEncoder, HwpxStyleStore};
use hwpforge_core::{Document, Section, Paragraph, PageSettings};
use hwpforge_core::run::Run;
use hwpforge_foundation::{CharShapeIndex, ParaShapeIndex};

// 새 문서 생성
let mut doc = Document::new();
doc.add_section(Section::with_paragraphs(
    vec![Paragraph::with_runs(
        vec![Run::text("안녕하세요, HwpForge!", CharShapeIndex::new(0))],
        ParaShapeIndex::new(0),
    )],
    PageSettings::a4(),
));

let validated = doc.validate().unwrap();
let style_store = HwpxStyleStore::with_default_fonts("함초롬바탕");
let image_store = Default::default();

let bytes = HwpxEncoder::encode(&validated, &style_store, &image_store).unwrap();
std::fs::write("output.hwpx", &bytes).unwrap();
}

HwpxStyleStore 생성 방법

with_default_fonts() — 간단한 기본 스타일

단일 글꼴 이름으로 빠르게 스타일 스토어를 생성합니다. 가장 간단한 방법입니다.

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxStyleStore;

let style_store = HwpxStyleStore::with_default_fonts("함초롬바탕");
}

from_registry() — Blueprint 템플릿에서 변환

커스텀 YAML 스타일 템플릿을 적용할 때 사용합니다. 자세한 내용은 스타일 템플릿 참조.

#![allow(unused)]
fn main() {
use hwpforge::blueprint::builtins::builtin_default;
use hwpforge::blueprint::registry::StyleRegistry;
use hwpforge::hwpx::HwpxRegistryBridge;

let template = builtin_default().unwrap();
let registry = StyleRegistry::from_template(&template).unwrap();
let bridge = HwpxRegistryBridge::from_registry(&registry).unwrap();
let style_store = bridge.style_store();
}

HwpxStyleStore::from_registry() 자체는 HWPX style table만 만듭니다.
Blueprint/Markdown 경로에서 만든 문서는 registry-local CharShapeIndex / ParaShapeIndex 를 들고 있으므로, encode 직전에는 HwpxRegistryBridge로 rebinding 해야 합니다.

라운드트립 예제 (decode → modify → encode)

기존 HWPX 파일을 읽어서 수정한 뒤 다시 저장하는 전형적인 패턴입니다.

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwpx::{HwpxDecoder, HwpxEncoder};
use hwpforge_core::run::Run;
use hwpforge_core::paragraph::Paragraph;
use hwpforge_foundation::{CharShapeIndex, ParaShapeIndex};

// 1. 기존 파일 디코딩
let mut result = HwpxDecoder::decode_file("original.hwpx").unwrap();

// 2. 문서 수정: 새 문단 추가
let new_para = Paragraph::with_runs(
    vec![Run::text("추가된 문단입니다.", CharShapeIndex::new(0))],
    ParaShapeIndex::new(0),
);
// Draft 상태이므로 sections 직접 접근 가능
result.document.sections_mut()[0].paragraphs.push(new_para);

// 3. 검증 후 인코딩
let validated = result.document.validate().unwrap();
let bytes = HwpxEncoder::encode(
    &validated,
    &result.style_store,
    &result.image_store,
).unwrap();

std::fs::write("modified.hwpx", &bytes).unwrap();
}

기존 텍스트 찾기 및 수정

특정 텍스트를 찾아 수정하려면 sections_mut()으로 가변 접근 후 RunContent::Text를 패턴 매칭합니다.

텍스트 치환 (find & replace)

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwpx::{HwpxDecoder, HwpxEncoder};
use hwpforge_core::run::RunContent;

// 1. 디코딩
let mut result = HwpxDecoder::decode_file("template.hwpx").unwrap();

// 2. 모든 섹션의 모든 문단을 순회하며 텍스트 치환
for section in result.document.sections_mut() {
    for paragraph in &mut section.paragraphs {
        for run in &mut paragraph.runs {
            if let RunContent::Text(ref mut text) = run.content {
                if text.contains("{{회사명}}") {
                    *text = text.replace("{{회사명}}", "한국테크");
                }
                if text.contains("{{날짜}}") {
                    *text = text.replace("{{날짜}}", "2026년 3월 11일");
                }
            }
        }
    }
}

// 3. 검증 후 저장
let validated = result.document.validate().unwrap();
let bytes = HwpxEncoder::encode(&validated, &result.style_store, &result.image_store).unwrap();
std::fs::write("output.hwpx", &bytes).unwrap();
}

재사용 가능한 치환 함수

#![allow(unused)]
fn main() {
use hwpforge_core::document::{Document, Draft};
use hwpforge_core::run::RunContent;

/// 문서 내 모든 텍스트에서 `from`을 `to`로 치환합니다.
/// 치환된 횟수를 반환합니다.
fn replace_text(doc: &mut Document<Draft>, from: &str, to: &str) -> usize {
    let mut count = 0;
    for section in doc.sections_mut() {
        for paragraph in &mut section.paragraphs {
            for run in &mut paragraph.runs {
                if let RunContent::Text(ref mut text) = run.content {
                    if text.contains(from) {
                        *text = text.replace(from, to);
                        count += 1;
                    }
                }
            }
        }
    }
    count
}
}

완전한 읽기 → 수정 → 저장 예제

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwpx::{HwpxDecoder, HwpxEncoder};
use hwpforge_core::run::{Run, RunContent};
use hwpforge_core::paragraph::Paragraph;
use hwpforge_foundation::{CharShapeIndex, ParaShapeIndex};

fn modify_document(
    input: &str,
    output: &str,
) -> Result<(), Box<dyn std::error::Error>> {
    // 읽기
    let mut result = HwpxDecoder::decode_file(input)
        .map_err(|e| format!("디코딩 실패: {e}"))?;

    let sections = result.document.sections_mut();

    // 기존 텍스트 수정
    for section in sections.iter_mut() {
        for paragraph in &mut section.paragraphs {
            for run in &mut paragraph.runs {
                if let RunContent::Text(ref mut text) = run.content {
                    *text = text.replace("초안", "최종본");
                }
            }
        }
    }

    // 새 문단 추가
    if let Some(first_section) = result.document.sections_mut().first_mut() {
        first_section.paragraphs.push(Paragraph::with_runs(
            vec![Run::text("— 이 문서는 자동으로 수정되었습니다.", CharShapeIndex::new(0))],
            ParaShapeIndex::new(0),
        ));
    }

    // 저장
    let validated = result.document.validate()
        .map_err(|e| format!("검증 실패: {e}"))?;
    let bytes = HwpxEncoder::encode(&validated, &result.style_store, &result.image_store)
        .map_err(|e| format!("인코딩 실패: {e}"))?;
    std::fs::write(output, &bytes)?;

    Ok(())
}
}

오류 처리

모든 함수는 HwpxResult<T>를 반환합니다. HwpxError는 HwpxErrorCode와 메시지를 포함합니다.

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;

match HwpxDecoder::decode_file("missing.hwpx") {
    Ok(_result) => println!("디코딩 성공"),
    Err(e) => eprintln!("디코딩 실패: {e}"),
}
}

엣지 케이스 및 주의사항

빈 문서

Document는 최소 1개의 섹션이 있어야 validate()를 통과합니다.

#![allow(unused)]
fn main() {
use hwpforge::core::{Document, Draft, PageSettings, Paragraph, Section};
use hwpforge::foundation::ParaShapeIndex;

let mut doc = Document::<Draft>::new();

// ❌ 빈 문서 — validate() 실패
// let validated = doc.validate();  // Err: 섹션 없음

// ✅ 빈 문단이라도 하나 추가
doc.add_section(Section::with_paragraphs(
    vec![Paragraph::new(ParaShapeIndex::new(0))],
    PageSettings::a4(),
));
let validated = doc.validate().unwrap();  // OK
}

한국어/특수 문자

HwpForge는 내부적으로 UTF-8을 사용합니다. 한국어, 이모지, 특수 기호를 포함한 모든 유니코드 문자를 지원합니다.

#![allow(unused)]
fn main() {
use hwpforge::core::run::Run;
use hwpforge::foundation::CharShapeIndex;

// 모두 정상 동작
let run1 = Run::text("한글 텍스트 테스트", CharShapeIndex::new(0));
let run2 = Run::text("특수문자: ©®™ §¶ ±×÷", CharShapeIndex::new(0));
let run3 = Run::text("수학 기호: α β γ δ ∑ ∫", CharShapeIndex::new(0));
}

스타일 스토어 선택

생성 방법용도특징
with_default_fonts("글꼴명")빠른 프로토타이핑한컴 Modern 22종 기본 스타일
from_registry(&registry)커스텀 템플릿 적용YAML로 정의한 스타일 사용
디코딩된 result.style_store기존 문서 수정원본 스타일 보존

메타데이터 (Metadata)

HwpForge의 모든 문서는 Metadata 구조체를 통해 제목, 작성자, 작성일 등의 메타데이터를 관리합니다.

Metadata 구조체

use std::collections::BTreeMap;

#[non_exhaustive]
pub struct Metadata {
    pub title: Option<String>,               // 문서 제목
    pub author: Option<String>,              // 작성자
    pub subject: Option<String>,             // 주제/설명
    pub description: Option<String>,         // 자유 서술 요약 (subject와 별개)
    pub last_saved_by: Option<String>,       // 마지막으로 저장한 사람 (author와 별개)
    pub keywords: Vec<String>,               // 검색 키워드
    pub created: Option<String>,             // 작성일 (ISO 8601, 예: "2026-03-06")
    pub modified: Option<String>,            // 수정일 (ISO 8601)
    pub extras: BTreeMap<String, String>,    // 아직 타입 필드로 승격되지 않은 <opf:meta> 등 원본 항목
}

모든 필드는 선택적입니다. Metadata::default()는 모든 필드가 비어 있는 상태를 반환합니다.

Metadata는 #[non_exhaustive]입니다 — 향후 버전에서 필드가 추가될 수 있으므로, 외부 크레이트에서는 ..Default::default()를 붙여도 구조체 리터럴로 생성할 수 없습니다. Metadata::new()에서 시작하는 빌더(with_title, with_author, with_subject, with_description, with_last_saved_by, with_keywords, with_created, with_modified, with_extra)를 사용하세요. 이미 만들어진 값의 필드는 doc.metadata_mut().title = ...처럼 직접 대입할 수 있습니다.

기존 HWPX 파일에서 메타데이터 읽기

HwpxDecoder로 HWPX 파일을 디코딩한 후 document.metadata()로 접근합니다.

use hwpforge::hwpx::HwpxDecoder;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let result = HwpxDecoder::decode_file("document.hwpx")?;
    let meta = result.document.metadata();

    // 개별 필드 접근
    if let Some(title) = &meta.title {
        println!("제목: {}", title);
    }
    if let Some(author) = &meta.author {
        println!("작성자: {}", author);
    }
    if let Some(created) = &meta.created {
        println!("작성일: {}", created);
    }
    if let Some(subject) = &meta.subject {
        println!("주제: {}", subject);
    }
    if !meta.keywords.is_empty() {
        println!("키워드: {}", meta.keywords.join(", "));
    }

    Ok(())
}

Markdown에서 메타데이터 설정

YAML Frontmatter로 메타데이터를 지정하면 MdDecoder가 자동으로 Metadata 필드에 매핑합니다.

#![allow(unused)]
fn main() {
use hwpforge::md::{MdDecoder, MdDocument};

let markdown = r#"---
title: 분기 보고서
author: 김철수
date: 2026-03-06
metadata:
  subject: 2026년 1분기 경영실적 보고
  keywords:
    - 분기실적
    - 경영보고
  modified: 2026-03-10
---

보고서 본문

내용이 여기에 들어갑니다.
"#;

let MdDocument { document, style_registry } = MdDecoder::decode_with_default(markdown).unwrap();

let meta = document.metadata();
assert_eq!(meta.title.as_deref(), Some("분기 보고서"));
assert_eq!(meta.author.as_deref(), Some("김철수"));
assert_eq!(meta.created.as_deref(), Some("2026-03-06"));
assert_eq!(meta.subject.as_deref(), Some("2026년 1분기 경영실적 보고"));
assert_eq!(meta.keywords, vec!["분기실적", "경영보고"]);
assert_eq!(meta.modified.as_deref(), Some("2026-03-10"));
}

Frontmatter 필드 매핑

최상위 필드:

YAML 필드Metadata 필드설명
titletitle문서 제목
authorauthor작성자
datecreated작성일 (ISO 8601)
template(없음)파싱되지만 스타일 선택에는 쓰이지 않습니다(inert)
metadata(중첩 맵)아래 하위 필드를 담는 컨테이너

metadata: 아래의 하위 필드:

하위 필드Metadata 필드설명
subjectsubject주제/설명
keywordskeywords검색 키워드 (배열)
modifiedmodified수정일 (ISO 8601)

subject/keywords/modified는 반드시 metadata: 아래에 중첩해야 합니다 — 최상위에 쓰면 조용히 무시됩니다.

프로그래밍으로 메타데이터 설정

Document<Draft> 상태에서 metadata_mut()으로 직접 설정할 수 있습니다.

#![allow(unused)]
fn main() {
use hwpforge::core::{Document, Draft, Metadata, PageSettings, Paragraph, Run, Section};
use hwpforge::foundation::{CharShapeIndex, ParaShapeIndex};

let mut doc = Document::<Draft>::new();

// 메타데이터 설정
doc.metadata_mut().title = Some("제안서".to_string());
doc.metadata_mut().author = Some("홍길동".to_string());
doc.metadata_mut().created = Some("2026-03-06".to_string());
doc.metadata_mut().subject = Some("신규 사업 제안".to_string());
doc.metadata_mut().keywords = vec!["사업".to_string(), "제안".to_string()];

// 또는 빌더로 Metadata를 만들어 한 번에 설정
let meta = Metadata::new()
    .with_title("제안서")
    .with_author("홍길동")
    .with_created("2026-03-06");
doc.set_metadata(meta);

// 섹션 추가 후 검증/인코딩
doc.add_section(Section::with_paragraphs(
    vec![Paragraph::with_runs(
        vec![Run::text("본문 내용", CharShapeIndex::new(0))],
        ParaShapeIndex::new(0),
    )],
    PageSettings::a4(),
));
let validated = doc.validate().unwrap();
}

CLI에서 메타데이터 확인

hwpforge inspect 명령으로 HWPX 파일의 메타데이터를 확인합니다.

# 사람이 읽기 좋은 출력
hwpforge inspect document.hwpx

# 출력 예시:
# Document: document.hwpx
#   Title:  분기 보고서
#   Author: 김철수
#   Sections: 1
#     [0] 2 paras (deep 2), 0 tables, 0 images, 0 charts | header=false footer=false pagenum=false
# JSON 출력 (AI 에이전트용)
hwpforge inspect document.hwpx --json

# 출력 예시:
# {
#   "status": "ok",
#   "metadata": {
#     "title": "분기 보고서",
#     "author": "김철수"
#   },
#   "sections": [
#     {
#       "index": 0,
#       "paragraphs": 2,
#       "deep_paragraphs": 2,
#       "tables": 0,
#       "images": 0,
#       "charts": 0,
#       "has_header": false,
#       "has_footer": false,
#       "has_page_number": false,
#       ...
#     }
#   ]
# }

JSON 라운드트립에서 메타데이터

to-json으로 내보내면 메타데이터가 JSON에 포함됩니다.

hwpforge to-json document.hwpx -o doc.json
{
  "document": {
    "sections": [...],
    "metadata": {
      "title": "분기 보고서",
      "author": "김철수",
      "subject": null,
      "description": null,
      "last_saved_by": null,
      "keywords": [],
      "created": "2026-03-06",
      "modified": null,
      "extras": {}
    }
  },
  "styles": {...}
}

AI 에이전트가 JSON에서 메타데이터를 수정한 후 from-json으로 HWPX를 재생성할 수 있습니다.

# JSON 편집 후 HWPX로 변환
hwpforge from-json doc.json -o updated.hwpx

MCP 도구에서 메타데이터 확인

hwpforge_inspect MCP 도구로 메타데이터를 포함한 문서 구조를 확인합니다.

{
  "tool": "hwpforge_inspect",
  "arguments": {
    "file_path": "/path/to/document.hwpx"
  }
}

현재 제한사항

  • HWPX 네이티브 메타데이터: 디코더는 한글 프로그램이 저장한 HWPX의 Contents/content.hpf (<opf:metadata>)를 읽어 제목(<opf:title>)과 <opf:meta name="..."> 항목 중 작성자(creator)·주제(subject)·설명(description)·마지막 저장자(lastsaveby)·작성일(CreatedDate)·수정일(ModifiedDate)·키워드(keyword, 세미콜론 구분)를 Metadata에 채웁니다. 아직 타입 필드가 없는 <opf:meta> 항목은 extras에 보존됩니다.
  • 타임스탬프 형식: created/modified는 Option<String> (ISO 8601 문자열)입니다. chrono 등 날짜 라이브러리와 연동 시 직접 파싱이 필요합니다.

Markdown에서 HWPX로

HwpForge는 Markdown을 HWPX로 변환하는 완전한 파이프라인을 제공합니다. LLM이 Markdown을 생성하면 HwpForge가 이를 한글 문서로 자동 변환합니다.

MD → Core → HWPX 파이프라인

Markdown 문자열
    |
    v (MdDecoder::decode)
Document<Draft> + StyleRegistry
    |
    v (doc.validate())
Document<Validated>
    |
    v (HwpxEncoder::encode)
HWPX 바이트 → .hwpx 파일

각 단계는 독립적이므로, 중간 Core DOM을 직접 조작하거나 검사할 수 있습니다.

MdDecoder::decode() 사용법

#![allow(unused)]
fn main() {
use hwpforge::md::{MdDecoder, MdDocument};

let markdown = r#"
---
title: 사업 제안서
author: 홍길동
date: 2026-03-06
---

개요

본 제안서는 신규 사업 기회를 설명합니다.

# 배경

시장 분석에 따르면 성장 가능성이 높습니다.
"#;

let MdDocument { document, style_registry } = MdDecoder::decode_with_default(markdown).unwrap();

println!("섹션 수: {}", document.sections().len());
}

MdDocument에는 document: Document<Draft>와 style_registry: StyleRegistry가 포함됩니다.

YAML Frontmatter

Markdown 파일 상단에 --- 블록으로 문서 메타데이터를 지정합니다.

---
title: 문서 제목          # Metadata.title
author: 작성자 이름        # Metadata.author
date: 2026-03-06          # Metadata.created (ISO 8601)
template: government      # 파싱되지만 스타일 선택에는 쓰이지 않습니다 (inert)
metadata:                 # 아래 하위 필드를 담는 중첩 맵
  subject: 신규 사업 제안   # Metadata.subject
  keywords:                # Metadata.keywords (YAML 배열)
    - 사업
    - 제안
  modified: 2026-03-10     # Metadata.modified (ISO 8601)
---

최상위 필드:

필드Metadata 필드설명
titletitle문서 제목
authorauthor작성자
datecreated작성일 (ISO 8601)
template(없음)파싱되지만 스타일 선택에는 쓰이지 않습니다(inert)
metadata(중첩 맵)아래 하위 필드를 담는 컨테이너

metadata: 아래의 하위 필드:

하위 필드Metadata 필드설명
subjectsubject주제/설명
keywordskeywords검색 키워드 (YAML 배열)
modifiedmodified수정일 (ISO 8601)

subject/keywords/modified는 반드시 metadata: 아래에 중첩해야 합니다 — 최상위에 쓰면 조용히 무시됩니다(Frontmatter 구조체가 이 키들을 최상위 필드로 갖지 않고, 모르는 키를 에러로 거부하지도 않기 때문입니다).

Frontmatter 없이도 디코딩이 가능하며, 메타데이터 필드는 빈 값으로 처리됩니다.

디코딩 후 메타데이터 확인

#![allow(unused)]
fn main() {
use hwpforge::md::{MdDecoder, MdDocument};

let markdown = "---\ntitle: 보고서\nauthor: 홍길동\ndate: 2026-03-06\n---\n\n# 본문\n";
let MdDocument { document, .. } = MdDecoder::decode_with_default(markdown).unwrap();

let meta = document.metadata();
assert_eq!(meta.title.as_deref(), Some("보고서"));
assert_eq!(meta.author.as_deref(), Some("홍길동"));
assert_eq!(meta.created.as_deref(), Some("2026-03-06"));
}

전체 메타데이터 필드와 프로그래밍 설정 방법은 메타데이터 가이드를 참고하세요.

섹션 마커

<!-- hwpforge:section --> 주석으로 HWPX 섹션을 분리합니다. 한 Markdown 파일에서 여러 섹션(페이지 설정이 다른 구역)을 만들 때 유용합니다.

# 1장 개요

첫 번째 섹션 내용.

<!-- hwpforge:section -->

# 2장 본론

두 번째 섹션 — 다른 페이지 설정 가능.

각주 (Footnote) / 미주 (Endnote)

GFM 각주 문법으로 각주와 미주를 표현합니다.

본문에 각주를 답니다.[^1] 미주도 답니다.[^e1]

[^1]: 각주 본문입니다.

[^e1]: 미주 본문입니다.
  • [^라벨]: 각주 — 관례적으로 숫자를 씁니다([^1]), 하지만 규칙 자체는 아래 미주 형태(e[0-9]+)가 아니면 어떤 라벨도 각주로 인정됩니다(예: [^note]도 유효한 각주 라벨입니다).
  • [^eN] (e + 숫자 1개 이상): 미주 — HwpForge dialect의 예약 네임스페이스입니다. 각주 의도로 [^e1]을 써도 미주로 정규화됩니다 (lossy 변환입니다).
  • 정의는 여러 문단으로 이어갈 수 있습니다. 첫 문단 다음에 빈 줄을 두고, 이어지는 문단의 모든 줄을 4-space 들여씁니다.
[^1]: 첫 번째 문단.

    두 번째 문단 (4-space 들여쓰기).

hwpforge to-md(기본 styled 모드)로 HWPX → Markdown 역방향 변환할 때도 각주는 [^N], 미주는 [^eN]으로 방출되어 왕복이 보존됩니다. lossy 모드는 대신 (footnote: ...)/(endnote: ...) 형태의 인라인 텍스트로 펼쳐서 출력하므로 이 왕복 규약의 대상이 아닙니다.

표 (Table)

GFM 표 문법을 지원합니다.

| 항목 | 값     |
| ---- | ------ |
| 이름 | 홍길동 |
| 부서 | 기획팀 |

표 셀 안에는 링크, 이미지, 각주/미주 참조 등 인라인 요소도 담을 수 있습니다.

이미지

![대체 텍스트](images/photo.png)
  • 지원 포맷: PNG, JPEG, GIF, BMP, WMF, EMF (SVG는 지원하지 않습니다) — 확장자가 아니라 실제 바이트를 스니핑해 포맷을 확인합니다.
  • 경로(상대·절대 모두)는 정규화한 뒤 Markdown 파일이 있는 디렉터리 하위에 있는지 검사합니다. ../로 그 디렉터리를 벗어나는 경로는 거부되지만, 그 디렉터리 안을 가리키는 절대 경로는 허용됩니다 — 거부되는 것은 “절대 경로“가 아니라 “벗어나는 경로“입니다.
  • data: URI(base64)도 지원합니다 — 네트워크 접근 없이 로컬에서 바로 디코드됩니다. 기준 디렉터리가 없는 인라인 텍스트/stdin 입력에서는 data: URI만 임베드할 수 있습니다.
  • http(s) 원격 URL은 네트워크 접근을 금지하는 정책상 거부됩니다.
  • 실패한 참조(파일 없음·경로 탈출·원격 URL·미지 포맷 등)는 경고와 함께 이미지 run이 드롭됩니다 — 결과 문서에 깨진 참조가 남지 않습니다.
  • 파일 크기 상한은 50MB입니다.

구분선과 섹션 구분자

Markdown의 ---(frontmatter 블록 밖에서 단독으로 쓴 thematic break)는 리터럴 "---" 문단으로 변환됩니다 — 페이지나 섹션을 나누는 구분자가 아닙니다. HWPX 섹션(다른 페이지 설정 구역)을 나누려면 반드시 <!-- hwpforge:section --> 주석을 사용하세요 (앞의 섹션 마커 참고).

H1-H6 → 개요 1-6 자동 매핑

Markdown 헤딩은 한글의 개요 스타일로 자동 변환됩니다.

Markdown한글 스타일
# H1개요 1 (style ID 2)
## H2개요 2 (style ID 3)
### H3개요 3 (style ID 4)
#### H4개요 4 (style ID 5)
##### H5개요 5 (style ID 6)
###### H6개요 6 (style ID 7)
일반 문단본문 (style ID 0)

MdEncoder — Core → Markdown

반대 방향(HWPX → Markdown) 변환도 지원합니다. 두 가지 모드가 있습니다.

#![allow(unused)]
fn main() {
use hwpforge::md::MdEncoder;
use hwpforge::hwpx::HwpxDecoder;

let result = HwpxDecoder::decode_file("document.hwpx").unwrap();
let validated = result.document.validate().unwrap();

// Lossy 모드: 읽기 좋은 GFM (표, 이미지 등 일부 정보 손실)
let gfm = MdEncoder::encode_lossy(&validated).unwrap();

// Lossless 모드: YAML frontmatter + HTML-like 마크업 (정보 보존)
let lossless = MdEncoder::encode_lossless(&validated).unwrap();
}
모드특징용도
encode_lossy읽기 좋은 GFM사람이 읽는 문서 미리보기
encode_lossless구조 완전 보존라운드트립, 백업

전체 파이프라인 예제 (MD string → HWPX file)

use hwpforge::md::{MdDecoder, MdDocument};
use hwpforge::hwpx::{HwpxEncoder, HwpxStyleStore};

fn markdown_to_hwpx(markdown: &str, output_path: &str) {
    // 1. Markdown 파싱 → Core DOM
    let MdDocument { document, .. } = MdDecoder::decode_with_default(markdown).unwrap();

    // 2. 문서 검증
    let validated = document.validate().unwrap();

    // 3. 한컴 기본 스타일 적용 후 HWPX 인코딩
    let style_store = HwpxStyleStore::with_default_fonts("함초롬바탕");
    let image_store = Default::default();
    let bytes = HwpxEncoder::encode(&validated, &style_store, &image_store).unwrap();

    // 4. 파일 저장
    std::fs::write(output_path, &bytes).unwrap();
    println!("저장 완료: {output_path}");
}

fn main() {
    let md = r#"
---
title: AI 활용 정책 제안서
author: 정책팀
date: 2026-03-06
---

제안 배경

인공지능 기술의 급속한 발전에 대응하여 정책 수립이 필요합니다.

# 현황 분석

국내외 AI 활용 사례를 분석하였습니다.

# 정책 방향

단계적 도입과 윤리적 기준 마련을 제안합니다.
"#;

    markdown_to_hwpx(md, "proposal.hwpx");
}

HWPX → Markdown 변환 (RAG/LLM 활용)

HWPX 문서를 Markdown으로 변환하면 LLM이나 RAG(Retrieval-Augmented Generation) 시스템에서 직접 활용할 수 있습니다.

의존성 설정

Cargo.toml에 md 기능을 활성화합니다:

[dependencies]
hwpforge = { version = "0.16", features = ["md"] }

Lossy vs Lossless 모드 선택

기준Lossy (encode_lossy)Lossless (encode_lossless)
출력 형식표준 GFM MarkdownYAML frontmatter + HTML 마크업
가독성높음 (사람/LLM 모두)낮음 (기계 파싱용)
정보 손실스타일/레이아웃 일부 손실구조 완전 보존
RAG 추천추천 — 청크 분할에 적합원본 복원이 필요할 때만
LLM 추천추천 — 토큰 효율적라운드트립 편집 시

RAG 시스템에서는 encode_lossy를 권장합니다. 표준 GFM으로 출력되어 청크 분할기(text splitter)와 호환성이 높고, 불필요한 마크업이 없어 토큰을 절약합니다.

완전한 HWPX → Markdown 예제 (에러 처리 포함)

use hwpforge::hwpx::HwpxDecoder;
use hwpforge::md::MdEncoder;
use std::path::Path;

fn hwpx_to_markdown(input_path: &str) -> Result<String, Box<dyn std::error::Error>> {
    // 1. 파일 존재 여부 확인
    let path = Path::new(input_path);
    if !path.exists() {
        return Err(format!("파일을 찾을 수 없습니다: {}", input_path).into());
    }

    // 2. HWPX 디코딩
    let result = HwpxDecoder::decode_file(input_path)
        .map_err(|e| format!("HWPX 디코딩 실패: {e}"))?;

    // 3. 메타데이터 확인 (선택)
    let meta = result.document.metadata();
    if let Some(title) = &meta.title {
        eprintln!("문서 제목: {}", title);
    }

    // 4. Draft → Validated 상태 전이
    let validated = result.document.validate()
        .map_err(|e| format!("문서 검증 실패: {e}"))?;

    // 5. Markdown 변환 (RAG용 lossy 모드)
    let markdown = MdEncoder::encode_lossy(&validated)
        .map_err(|e| format!("Markdown 인코딩 실패: {e}"))?;

    Ok(markdown)
}

fn main() {
    match hwpx_to_markdown("document.hwpx") {
        Ok(md) => {
            std::fs::write("output.md", &md).expect("파일 저장 실패");
            println!("변환 완료: {} bytes", md.len());
        }
        Err(e) => eprintln!("오류: {e}"),
    }
}

대량 파일 변환

여러 HWPX 파일을 Markdown으로 일괄 변환합니다:

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;
use hwpforge::md::MdEncoder;
use std::path::Path;

fn batch_convert(input_dir: &str, output_dir: &str) -> Result<usize, Box<dyn std::error::Error>> {
    std::fs::create_dir_all(output_dir)?;
    let mut count = 0;

    for entry in std::fs::read_dir(input_dir)? {
        let entry = entry?;
        let path = entry.path();

        if path.extension().is_some_and(|ext| ext == "hwpx") {
            let result = HwpxDecoder::decode_file(&path)?;
            let validated = result.document.validate()?;
            let markdown = MdEncoder::encode_lossy(&validated)?;

            let out_name = path.file_stem().unwrap().to_string_lossy();
            let out_path = Path::new(output_dir).join(format!("{}.md", out_name));
            std::fs::write(&out_path, &markdown)?;

            eprintln!("변환: {} → {}", path.display(), out_path.display());
            count += 1;
        }
    }

    Ok(count)
}
}

CLI로 변환

# Markdown → HWPX
hwpforge convert report.md -o report.hwpx

# HWPX → Markdown
hwpforge to-md report.hwpx -o report.md

# 변환 모드 선택 (기본값: styled)
hwpforge to-md report.hwpx -o report.md --mode lossy
hwpforge to-md report.hwpx -o report.md --mode lossless

# HWPX 구조 확인 후 JSON으로 추출 (Markdown 변환 대안)
hwpforge inspect document.hwpx --json
hwpforge to-json document.hwpx -o document.json

참고: CLI의 convert 명령은 Markdown → HWPX 방향만 지원합니다. HWPX → Markdown 변환은 hwpforge to-md 명령(모드 styled/lossy/lossless) 또는 Rust API(MdEncoder)를 사용하세요.

스타일 템플릿 (YAML)

Blueprint 개념: 구조와 스타일 분리

HwpForge는 HTML+CSS와 동일한 철학으로 **구조(Core)**와 **스타일(Blueprint)**을 분리합니다.

Core (Document, Section, Paragraph, Run)
    = HTML — "무엇이 있는가"

Blueprint (Template, StyleRegistry, CharShape, ParaShape)
    = CSS  — "어떻게 보이는가"

Core의 문단과 런은 스타일 인덱스(ParaShapeIndex, CharShapeIndex)만 참조합니다. 실제 폰트 이름이나 크기는 Blueprint에 정의됩니다. 덕분에 동일한 문서 구조에 다른 템플릿을 적용해 전혀 다른 외관의 HWPX를 생성할 수 있습니다.

Template YAML 구조

meta:
  name: my-template
  version: "1.0"
  description: "커스텀 스타일 템플릿"

styles:
  body:
    font: "한컴바탕"
    size: 10pt
    line_spacing: 160%
    alignment: justify

  heading1:
    inherits: body        # body에서 상속
    font: "한컴고딕"
    size: 16pt
    bold: true

  heading2:
    inherits: heading1
    size: 14pt

상속 (Inheritance)

inherits 키로 다른 스타일을 상속받습니다. 상속 체인은 DFS로 해결되며, 자식 스타일의 값이 부모를 덮어씁니다. Option 필드(PartialCharShape, PartialParaShape)를 병합하는 two-type 패턴으로 구현됩니다.

StyleRegistry: from_template() 사용법

#![allow(unused)]
fn main() {
use hwpforge_blueprint::template::Template;
use hwpforge_blueprint::registry::StyleRegistry;

let yaml = r#"
meta:
  name: custom
  version: "1.0"
styles:
  body:
    font: "나눔명조"
    size: 11pt
"#;

// YAML → Template → StyleRegistry
let template = Template::from_yaml(yaml).unwrap();
let registry = StyleRegistry::from_template(&template).unwrap();

// 인덱스 기반 접근 (브랜드 타입으로 혼용 방지)
let body_entry = registry.get_style("body").unwrap();
let char_shape = registry.char_shape(body_entry.char_shape_id).unwrap();
println!("폰트: {}", char_shape.font);       // "나눔명조"
println!("크기: {:?}", char_shape.size);     // HwpUnit
}

내장 템플릿: builtin_default()

별도 YAML 없이 즉시 사용 가능한 기본 템플릿입니다.

#![allow(unused)]
fn main() {
use hwpforge_blueprint::builtins::builtin_default;
use hwpforge_blueprint::registry::StyleRegistry;

let template = builtin_default().unwrap();
assert_eq!(template.meta.name, "default");

let registry = StyleRegistry::from_template(&template).unwrap();
let body = registry.get_style("body").unwrap();
let cs = registry.char_shape(body.char_shape_id).unwrap();
assert_eq!(cs.font, "한컴바탕");
}

HwpxRegistryBridge 변환: from_registry()

Blueprint의 StyleRegistry를 HWPX 인코더 경계에서 안전하게 쓰기 위한 bridge를 만듭니다. 이 bridge는 두 가지를 함께 맡습니다.

  • HwpxStyleStore 생성
  • registry-local CharShapeIndex / ParaShapeIndex를 store-local HWPX id로 rebinding
#![allow(unused)]
fn main() {
use hwpforge_blueprint::builtins::builtin_default;
use hwpforge_blueprint::registry::StyleRegistry;
use hwpforge_smithy_hwpx::HwpxRegistryBridge;

let template = builtin_default().unwrap();
let registry = StyleRegistry::from_template(&template).unwrap();

// Blueprint StyleRegistry → HWPX encode bridge
let bridge = HwpxRegistryBridge::from_registry(&registry).unwrap();
let style_store = bridge.style_store();
}

한컴 스타일셋: Classic / Modern / Latest

한글 프로그램은 버전에 따라 다른 기본 스타일 구성을 사용합니다.

스타일셋스타일 수설명
Classic18개한글 2014–2020
Modern22개기본값 (한글 2022 이후)
Latest23개한글 2025 이후

Modern은 개요 8/9/10을 스타일 ID 9-11에 삽입하므로 인덱스가 Classic과 다릅니다.

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwpx::{HwpxStyleStore, HancomStyleSet};

// 기본값 (간단한 방법)
let modern = HwpxStyleStore::with_default_fonts("함초롬바탕");

// 특정 스타일셋 지정
// from_registry_with()로 커스텀 레지스트리 + 스타일셋 조합 가능
}

예제: 커스텀 스타일로 문서 생성

#![allow(unused)]
fn main() {
use hwpforge_blueprint::template::Template;
use hwpforge_blueprint::registry::StyleRegistry;
use hwpforge_smithy_hwpx::{HwpxEncoder, HwpxRegistryBridge};
use hwpforge_core::{Document, Section, Paragraph, PageSettings};
use hwpforge_core::run::Run;
use hwpforge_foundation::{CharShapeIndex, ParaShapeIndex};

let yaml = r#"
meta:
  name: report
  version: "1.0"
styles:
  body:
    font: "맑은 고딕"
    size: 10pt
    line_spacing: 150%
  title:
    inherits: body
    font: "맑은 고딕"
    size: 20pt
    bold: true
    alignment: center
"#;

// 스타일 빌드
let template = Template::from_yaml(yaml).unwrap();
let registry = StyleRegistry::from_template(&template).unwrap();
let bridge = HwpxRegistryBridge::from_registry(&registry).unwrap();

// 문서 구성 (스타일 인덱스는 레지스트리에서 조회)
let mut doc = Document::new();
doc.add_section(Section::with_paragraphs(
    vec![
        // 제목 문단 (ParaShapeIndex 0 = title)
        Paragraph::with_runs(
            vec![Run::text("분기 보고서", CharShapeIndex::new(0))],
            ParaShapeIndex::new(0),
        ),
        // 본문 문단 (ParaShapeIndex 1 = body)
        Paragraph::with_runs(
            vec![Run::text("1분기 실적은 목표를 초과 달성했습니다.", CharShapeIndex::new(1))],
            ParaShapeIndex::new(1),
        ),
    ],
    PageSettings::a4(),
));

let rebound = bridge.rebind_draft_document(doc).unwrap();
let validated = rebound.validate().unwrap();
let image_store = Default::default();
let bytes = HwpxEncoder::encode(&validated, bridge.style_store(), &image_store).unwrap();
std::fs::write("report.hwpx", &bytes).unwrap();
}

차트 생성

HwpForge는 OOXML 차트 형식(xmlns:c)을 사용해 18종의 차트를 HWPX 문서에 삽입할 수 있습니다.

지원 차트 종류 (18종)

변형설명
Bar가로 막대 차트
Column세로 막대 차트
Bar3D / Column3D3D 막대/세로 막대
Line / Line3D꺾은선 / 3D 꺾은선
Pie / Pie3D원형 / 3D 원형
Doughnut도넛 차트
OfPie원형-of-원형 / 막대-of-원형
Area / Area3D영역 / 3D 영역
Scatter분산형 (XY)
Bubble버블 차트
Radar방사형 차트
Surface / Surface3D표면 / 3D 표면
Stock주식 차트 (HLC/OHLC/VHLC/VOHLC)

Control::Chart 생성 방법

차트는 Control::Chart 변형으로 표현됩니다. Run::control()로 런에 삽입하고, 그 런을 문단에 넣습니다.

#![allow(unused)]
fn main() {
use hwpforge_core::control::Control;
use hwpforge_core::chart::{ChartType, ChartData, ChartGrouping, LegendPosition};
use hwpforge_foundation::HwpUnit;

let chart = Control::Chart {
    chart_type: ChartType::Column,
    data: ChartData::category(
        &["1월", "2월", "3월", "4월"],
        &[("매출", &[1200.0, 1500.0, 1350.0, 1800.0])],
    ),
    title: Some("월별 매출".to_string()),
    legend: LegendPosition::Bottom,
    grouping: ChartGrouping::Clustered,
    width: HwpUnit::from_mm(120.0).unwrap(),
    height: HwpUnit::from_mm(80.0).unwrap(),
    bar_shape: None,
    explosion: None,
    of_pie_type: None,
    radar_style: None,
    wireframe: None,
    bubble_3d: None,
    scatter_style: None,
    show_markers: None,
    stock_variant: None,
};
}

Control::Chart에는 위 7개 외에 차트 종류별 옵션 필드 9개(bar_shape, explosion, of_pie_type, radar_style, wireframe, bubble_3d, scatter_style, show_markers, stock_variant)가 있습니다. 모두 Option이며 None이면 해당 차트의 기본 모양입니다. 제목·크기 등을 바꿀 필요가 없다면 기본값 생성자 Control::chart(chart_type, data)가 더 간단합니다(너비 약 114mm, 높이 약 66mm, 제목 없음, 범례 오른쪽, 그룹 방식 ChartGrouping::Clustered(기본값, 계열을 나란히 배치)).

ChartData: Category vs Xy 방식

Category 방식 (막대, 꺾은선, 원형, 영역, 방사형 등)

카테고리 레이블(X축)과 여러 시리즈로 구성됩니다. 대부분의 차트 종류에 사용합니다.

#![allow(unused)]
fn main() {
use hwpforge_core::chart::ChartData;

// 편의 생성자: cats 슬라이스 + (이름, 값 슬라이스) 튜플 배열
let data = ChartData::category(
    &["1분기", "2분기", "3분기", "4분기"],
    &[
        ("매출액", &[4200.0, 5100.0, 4800.0, 6200.0]),
        ("비용", &[3100.0, 3400.0, 3200.0, 3900.0]),
    ],
);
}

Xy 방식 (분산형, 버블)

X값과 Y값 쌍으로 구성됩니다. 두 변수 간의 관계를 나타낼 때 사용합니다.

#![allow(unused)]
fn main() {
use hwpforge_core::chart::ChartData;

// (이름, x값 슬라이스, y값 슬라이스) 튜플 배열
let data = ChartData::xy(&[
    ("데이터셋 A", &[1.0, 2.0, 3.0, 4.0], &[2.1, 3.9, 6.2, 7.8]),
    ("데이터셋 B", &[1.0, 2.0, 3.0, 4.0], &[1.5, 3.0, 5.0, 6.5]),
]);
}

ChartSeries, XySeries 구조

시리즈를 직접 구성할 때는 구조체를 사용합니다.

#![allow(unused)]
fn main() {
use hwpforge_core::chart::{ChartData, ChartSeries, XySeries};

// Category용 시리즈
let series = ChartSeries {
    name: "판매량".to_string(),
    values: vec![100.0, 150.0, 200.0],
};

let data = ChartData::Category {
    categories: vec!["A".to_string(), "B".to_string(), "C".to_string()],
    series: vec![series],
};

// XY용 시리즈
let xy_series = XySeries {
    name: "측정값".to_string(),
    x_values: vec![0.0, 1.0, 2.0],
    y_values: vec![0.0, 1.0, 4.0],
};
}

차트를 문단에 삽입하는 패턴

차트 Control을 Run::control()로 감싼 뒤, Paragraph::with_runs()에 포함시킵니다.

#![allow(unused)]
fn main() {
use hwpforge_core::control::Control;
use hwpforge_core::chart::{ChartType, ChartData};
use hwpforge_core::run::Run;
use hwpforge_core::paragraph::Paragraph;
use hwpforge_foundation::{CharShapeIndex, ParaShapeIndex};

let chart_control = Control::chart(
    ChartType::Column,
    ChartData::category(
        &["A", "B", "C"],
        &[("값", &[10.0, 20.0, 30.0])],
    ),
);

let para = Paragraph::with_runs(
    vec![Run::control(chart_control, CharShapeIndex::new(0))],
    ParaShapeIndex::new(0),
);
}

예제: 막대 차트 (Column)

#![allow(unused)]
fn main() {
use hwpforge_core::control::Control;
use hwpforge_core::chart::{ChartType, ChartData, ChartGrouping, LegendPosition};
use hwpforge_core::run::Run;
use hwpforge_core::paragraph::Paragraph;
use hwpforge_core::{Document, Section, PageSettings};
use hwpforge_smithy_hwpx::{HwpxEncoder, HwpxStyleStore};
use hwpforge_foundation::{CharShapeIndex, ParaShapeIndex, HwpUnit};

let data = ChartData::category(
    &["2022", "2023", "2024", "2025"],
    &[
        ("국내 매출", &[3200.0, 4100.0, 5300.0, 6800.0]),
        ("해외 매출", &[1100.0, 1800.0, 2700.0, 3900.0]),
    ],
);

let chart = Control::Chart {
    chart_type: ChartType::Column,
    data,
    title: Some("연도별 매출 현황 (단위: 백만원)".to_string()),
    legend: LegendPosition::Bottom,
    grouping: ChartGrouping::Clustered,
    width: HwpUnit::from_mm(140.0).unwrap(),
    height: HwpUnit::from_mm(90.0).unwrap(),
    bar_shape: None,
    explosion: None,
    of_pie_type: None,
    radar_style: None,
    wireframe: None,
    bubble_3d: None,
    scatter_style: None,
    show_markers: None,
    stock_variant: None,
};

let mut doc = Document::new();
doc.add_section(Section::with_paragraphs(
    vec![Paragraph::with_runs(
        vec![Run::control(chart, CharShapeIndex::new(0))],
        ParaShapeIndex::new(0),
    )],
    PageSettings::a4(),
));

let validated = doc.validate().unwrap();
let bytes = HwpxEncoder::encode(
    &validated,
    &HwpxStyleStore::with_default_fonts("함초롬바탕"),
    &Default::default(),
).unwrap();
std::fs::write("bar_chart.hwpx", &bytes).unwrap();
}

예제: 원형 차트 (Pie)

#![allow(unused)]
fn main() {
use hwpforge_core::control::Control;
use hwpforge_core::chart::{ChartType, ChartData, ChartGrouping, LegendPosition};
use hwpforge_foundation::{HwpUnit};

// 원형 차트는 단일 시리즈 사용
let chart = Control::Chart {
    chart_type: ChartType::Pie,
    data: ChartData::category(
        &["서울", "경기", "부산", "기타"],
        &[("비율", &[38.5, 25.2, 12.8, 23.5])],
    ),
    title: Some("지역별 매출 비중".to_string()),
    legend: LegendPosition::Right,
    grouping: ChartGrouping::Standard, // Pie는 Standard 사용
    width: HwpUnit::from_mm(100.0).unwrap(),
    height: HwpUnit::from_mm(80.0).unwrap(),
    bar_shape: None,
    explosion: None,
    of_pie_type: None,
    radar_style: None,
    wireframe: None,
    bubble_3d: None,
    scatter_style: None,
    show_markers: None,
    stock_variant: None,
};
}

주의: 차트 XML은 ZIP에 포함되지만 content.hpf 매니페스트에는 등록하지 않습니다. 매니페스트에 등록하면 한글이 크래시합니다.

텍스트 추출 (Text Extraction)

HwpForge의 Core DOM을 활용하여 문서에서 텍스트를 추출하고, 문서 구조(섹션, 문단, 표, 각주 등)를 보존하는 방법을 설명합니다.

포맷 지원 현황: 현재 HWPX(.hwpx)와 Markdown(.md)는 이 가이드의 예제대로 바로 텍스트 추출할 수 있습니다. 레거시 HWP5(.hwp)는 전용 crate/CLI 경로가 이미 존재하지만, top-level guide는 아직 HWPX/Markdown 중심으로 설명합니다. 자세한 내용은 이중 포맷 파이프라인을 참고하세요.

문서 구조 개요

HwpForge 문서는 다음과 같은 트리 구조를 가집니다:

Document
├── Metadata (title, author, created, ...)
├── Section 0
│   ├── PageSettings (용지 크기, 여백)
│   ├── headers / footers (Vec, 비어 있을 수 있음) / PageNumber (선택)
│   ├── Paragraph 0
│   │   ├── para_shape (문단 스타일 인덱스)
│   │   └── Run[]
│   │       ├── Run { content: Text("본문 텍스트"), char_shape }
│       ├── Run { content: InlineText(...), char_shape }   (탭 등 속성이 있는 텍스트)
│   │       ├── Run { content: Table(...), char_shape }
│   │       ├── Run { content: Image(...), char_shape }
│   │       └── Run { content: Control(Footnote/TextBox/...), char_shape }
│   ├── Paragraph 1
│   │   └── ...
│   └── ...
├── Section 1
│   └── ...
└── ...

RunContent는 #[non_exhaustive]라서 match에는 마지막에 _ => {} 팔이 필요합니다. 텍스트만 필요하다면 Text와 InlineText를 모두 처리하는 run.content.plain_text()(Option<Cow<str>>) 또는 문단 단위의 paragraph.text_content()를 쓰는 것이 안전합니다. RunContent::Text만 읽으면 InlineText 런의 텍스트가 조용히 빠집니다.

핵심 타입:

타입설명
RunContent::Text(String)일반 텍스트
RunContent::InlineText(InlineText)탭처럼 속성이 있는 텍스트 (plain_text()로 문자열화)
RunContent::Table(Box<Table>)인라인 표
RunContent::Image(Image)인라인 이미지
RunContent::Control(Box<Control>)컨트롤 (글상자, 하이퍼링크, 각주, 도형 등)

기본 텍스트 추출

가장 간단한 패턴: 모든 섹션의 모든 문단에서 텍스트만 추출합니다.

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;

let result = HwpxDecoder::decode_file("document.hwpx").unwrap();
let doc = &result.document;

for section in doc.sections() {
    for paragraph in &section.paragraphs {
        // Text 와 InlineText 런의 텍스트를 이어 붙인 문자열
        println!("{}", paragraph.text_content());
    }
}
}

구조 보존 텍스트 추출

문서 구조(섹션, 문단, 표, 각주 등)를 보존하면서 텍스트를 추출합니다.

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;
use hwpforge::core::run::RunContent;
use hwpforge::core::control::Control;
use hwpforge::core::paragraph::Paragraph;

let result = HwpxDecoder::decode_file("document.hwpx").unwrap();
let doc = &result.document;

// 메타데이터 출력
let meta = doc.metadata();
if let Some(title) = &meta.title {
    println!("=== {} ===", title);
}

for (sec_idx, section) in doc.sections().iter().enumerate() {
    println!("\n--- 섹션 {} ---", sec_idx + 1);

    // 머리글 텍스트 (적용 쪽 종류별로 여러 개일 수 있음)
    for header in &section.headers {
        print!("[머리글] ");
        extract_paragraphs(&header.paragraphs);
        println!();
    }

    // 본문 문단
    for paragraph in &section.paragraphs {
        extract_paragraph(paragraph, 0);
    }

    // 바닥글 텍스트
    for footer in &section.footers {
        print!("[바닥글] ");
        extract_paragraphs(&footer.paragraphs);
        println!();
    }
}

/// 단일 문단에서 텍스트 추출 (들여쓰기 레벨 지원)
fn extract_paragraph(para: &Paragraph, indent: usize) {
    let prefix = "  ".repeat(indent);
    print!("{}", prefix);

    for run in &para.runs {
        match &run.content {
            RunContent::Text(text) => print!("{}", text),
            RunContent::InlineText(inline) => print!("{}", inline.plain_text()),
            RunContent::Table(table) => {
                println!("\n{}[표 {}x{}]", prefix, table.row_count(), table.col_count());
                for (r, row) in table.rows.iter().enumerate() {
                    for (c, cell) in row.cells.iter().enumerate() {
                        print!("{}  [{},{}] ", prefix, r, c);
                        extract_paragraphs(&cell.paragraphs);
                    }
                }
            }
            RunContent::Image(img) => {
                print!("[이미지: {}]", img.path);
            }
            RunContent::Control(ctrl) => {
                extract_control(ctrl);
            }
            // RunContent 는 #[non_exhaustive] — 이후 추가될 변형은 건너뜀
            _ => {}
        }
    }
    println!();
}

/// 컨트롤 요소에서 텍스트 추출
fn extract_control(ctrl: &Control) {
    match ctrl {
        Control::TextBox { paragraphs, .. } => {
            print!("[글상자] ");
            extract_paragraphs(paragraphs);
        }
        Control::Hyperlink { text, url, .. } => {
            print!("[링크: {} → {}]", text, url);
        }
        Control::Footnote { paragraphs, .. } => {
            print!("[각주: ");
            extract_paragraphs(paragraphs);
            print!("]");
        }
        Control::Endnote { paragraphs, .. } => {
            print!("[미주: ");
            extract_paragraphs(paragraphs);
            print!("]");
        }
        // 도형 (Line, Ellipse, Polygon 등)은 텍스트 없음 — 건너뜀
        _ => {}
    }
}

/// 문단 목록에서 텍스트 추출 (헬퍼)
fn extract_paragraphs(paragraphs: &[Paragraph]) {
    for para in paragraphs {
        print!("{}", para.text_content());
    }
}
}

콘텐츠 요약 (빠른 분석)

문서의 구조적 특성을 빠르게 파악하려면 content_counts()를 사용합니다.

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;

let result = HwpxDecoder::decode_file("document.hwpx").unwrap();

for (i, section) in result.document.sections().iter().enumerate() {
    let counts = section.content_counts();
    println!(
        "섹션 {}: {} 문단, {} 표, {} 이미지, {} 차트",
        i, section.paragraphs.len(), counts.tables, counts.images, counts.charts
    );
    println!(
        "  머리글={}개 바닥글={}개 쪽번호={}",
        section.headers.len(),
        section.footers.len(),
        section.page_number.is_some()
    );
}
}

CLI로 텍스트 추출

inspect — 구조 요약

hwpforge inspect document.hwpx --json

to-json — 전체 DOM을 JSON으로 내보내기

JSON 출력에는 모든 텍스트와 구조 정보가 포함됩니다.

hwpforge to-json document.hwpx -o doc.json

Markdown 변환 — 읽기 쉬운 텍스트 추출

Rust API를 통해 HWPX를 Markdown으로 변환하면 구조를 보존한 텍스트를 얻을 수 있습니다.

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;
use hwpforge::md::MdEncoder;

let result = HwpxDecoder::decode_file("document.hwpx").unwrap();
let validated = result.document.validate().unwrap();

// 사람이 읽기 좋은 GFM (헤딩, 표, 목록 구조 보존)
let markdown = MdEncoder::encode_lossy(&validated).unwrap();
println!("{}", markdown);
}

이 방법은 문서 구조(헤딩 계층, 표, 목록, 인용 등)를 Markdown 형식으로 자연스럽게 보존합니다.

레거시 HWP5 파일 처리

레거시 HWP5(.hwp) 파일은 현재도 다룰 수 있습니다. 다만 public guide의 중심 경로는 아직 HWPX/Markdown 쪽입니다.

현재 선택지는 이렇습니다.

  1. CLI workflow 사용: convert-hwp5, audit-hwp5, census-hwp5
  2. 전용 crate 사용: hwpforge-smithy-hwp5의 Hwp5Decoder
  3. HWPX로 재출력 후 기존 guide 재사용: 변환 결과를 HWPX guide와 같은 방식으로 처리

전용 crate 경로 예시는 다음과 같습니다.

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwp5::Hwp5Decoder;

let result = Hwp5Decoder::decode_file("legacy.hwp").unwrap();
let doc = &result.document;

for section in doc.sections() {
    for paragraph in &section.paragraphs {
        println!("{}", paragraph.text_content());
    }
}
}

주의:

  • HWP5 경로는 warning-first가 기본입니다.
  • visual parity나 layout fidelity는 HWPX path보다 더 까다롭습니다.
  • stable top-level facade는 여전히 HWPX/Markdown 중심이므로, HWP5는 전용 crate 또는 CLI를 우선 보십시오.

HWP5와 HWPX: 이중 포맷 파이프라인

HwpForge는 한국의 두 가지 주요 문서 포맷 — 바이너리 OLE 기반 HWP5와 XML 기반 HWPX — 을 하나의 통합 파이프라인으로 처리할 수 있도록 설계되었습니다.

포맷 비교

특성HWP5 (.hwp)HWPX (.hwpx)
컨테이너OLE2/CFB (Compound File Binary)ZIP
내부 데이터바이너리 레코드 스트림XML 파일 (KS X 6101 OWPML)
표준한컴 독자 포맷 (공개 스펙)국가 표준 KS X 6101
역사1990년대~현재 (레거시)2014년~ (현대)
파일 시그니처D0 CF 11 E0 A1 B1 1A E1 (OLE)50 4B 03 04 (ZIP/PK)
스트림 구조FileHeader, DocInfo, BodyText/Section0 등mimetype, Contents/header.xml, Contents/section0.xml 등
압축zlib (스트림 단위)ZIP deflate (파일 단위)
암호화지원 (스트림 암호화)지원 (ZIP 암호화)
한글 호환성한글 97~최신한글 2014~최신

Core DOM: 포맷 독립 중간 표현 (IR)

HwpForge의 핵심 설계 원칙은 Core DOM이 포맷에 독립적이라는 것입니다. Document<Draft>는 HWP5든 HWPX든 Markdown이든 동일한 구조체로 표현됩니다.

┌─────────────┐     ┌─────────────┐     ┌─────────────┐
│  HWP5 파일  │     │  HWPX 파일  │     │  Markdown   │
│ (OLE/CFB)   │     │ (ZIP/XML)   │     │ (GFM+YAML)  │
└──────┬──────┘     └──────┬──────┘     └──────┬──────┘
       │ decode            │ decode            │ decode
       ▼                   ▼                   ▼
┌──────────────────────────────────────────────────────┐
│              Document<Draft>  (Core DOM)              │
│  ┌──────────────────────────────────────────────┐    │
│  │ Sections → Paragraphs → Runs → Text/Control  │    │
│  │ + Metadata (title, author, created, ...)      │    │
│  │ + Tables, Images, Shapes, Charts, ...         │    │
│  └──────────────────────────────────────────────┘    │
│              포맷 독립 중간 표현 (IR)                  │
└──────────┬───────────────┬───────────────┬───────────┘
           │ encode        │ encode        │ encode
           ▼               ▼               ▼
    ┌──────────┐    ┌──────────┐    ┌──────────┐
    │ HWP5     │    │ HWPX     │    │ Markdown │
    │ 미지원   │    │ ✅ 구현   │    │ ✅ 구현   │
    └──────────┘    └──────────┘    └──────────┘

이 설계 덕분에:

  • 하나의 문서 모델로 모든 포맷을 처리합니다
  • 포맷 간 변환이 Core DOM을 경유하여 자연스럽게 이루어집니다
  • 새 포맷 추가 시 기존 코드 수정 없이 Smithy 크레이트만 추가하면 됩니다
  • 비즈니스 로직은 Core DOM에만 의존하므로 포맷 변경에 영향을 받지 않습니다

포맷 감지

파일의 첫 바이트(매직 바이트)로 포맷을 판별합니다.

#![allow(unused)]
fn main() {
/// 파일 포맷 감지
enum DocumentFormat {
    Hwp5,     // OLE2/CFB 바이너리
    Hwpx,     // ZIP + XML
    Markdown, // 텍스트
    Unknown,
}

fn detect_format(bytes: &[u8]) -> DocumentFormat {
    if bytes.len() < 4 {
        return DocumentFormat::Unknown;
    }

    // OLE2 Compound File Binary: D0 CF 11 E0
    if bytes.starts_with(&[0xD0, 0xCF, 0x11, 0xE0]) {
        return DocumentFormat::Hwp5;
    }

    // ZIP (PK\x03\x04)
    if bytes.starts_with(&[0x50, 0x4B, 0x03, 0x04]) {
        return DocumentFormat::Hwpx;
    }

    // UTF-8 텍스트로 시작하면 Markdown 후보
    if std::str::from_utf8(bytes).is_ok() {
        return DocumentFormat::Markdown;
    }

    DocumentFormat::Unknown
}
}

포맷 독립 문서 처리

Core DOM을 활용하면 입력 포맷에 관계없이 동일한 코드로 문서를 처리할 수 있습니다.

현재 지원되는 파이프라인

#![allow(unused)]
fn main() {
use hwpforge::hwpx::{HwpxDecoder, HwpxEncoder, HwpxRegistryBridge};
use hwpforge::md::{MdDecoder, MdDocument, MdEncoder};
use hwpforge::core::{Document, Draft, ImageStore};

// === 1. HWPX → Core DOM ===
let hwpx_result = HwpxDecoder::decode_file("input.hwpx").unwrap();
let doc_from_hwpx: Document<Draft> = hwpx_result.document;

// === 2. Markdown → Core DOM ===
let markdown = "# 제목\n\n본문 내용입니다.";
let MdDocument { document: doc_from_md, style_registry } =
    MdDecoder::decode_with_default(markdown).unwrap();

// === 3. 포맷 독립 처리 (어느 소스에서 왔든 동일) ===
fn process_document(doc: &Document<Draft>) {
    // 메타데이터 접근
    let meta = doc.metadata();
    println!("제목: {:?}", meta.title);
    println!("작성자: {:?}", meta.author);

    // 섹션/문단 순회
    for section in doc.sections() {
        println!("문단 수: {}", section.paragraphs.len());
        let counts = section.content_counts();
        println!("표: {}, 이미지: {}", counts.tables, counts.images);
    }
}

process_document(&doc_from_hwpx);
process_document(&doc_from_md);

// === 4. Core DOM → 다른 포맷으로 출력 ===
// HWPX로 저장
let bridge = HwpxRegistryBridge::from_registry(&style_registry).unwrap();
let rebound = bridge.rebind_draft_document(doc_from_md).unwrap();
let validated = rebound.validate().unwrap();
let bytes = HwpxEncoder::encode(&validated, bridge.style_store(), &ImageStore::new()).unwrap();
std::fs::write("output.hwpx", &bytes).unwrap();

// Markdown으로 저장
let markdown_out = MdEncoder::encode_lossy(&validated).unwrap();
std::fs::write("output.md", &markdown_out).unwrap();
}

현재: HWP5 전용 crate / CLI 경로

#![allow(unused)]
fn main() {
use hwpforge_smithy_hwp5::Hwp5Decoder;
use hwpforge::hwpx::{HwpxEncoder, HwpxStyleStore};
use hwpforge::core::{Document, Draft, ImageStore};

let hwp5_result = Hwp5Decoder::decode_file("legacy.hwp").unwrap();
let doc: Document<Draft> = hwp5_result.document;

// Core DOM을 경유하여 HWP5 → HWPX 변환
let validated = doc.validate().unwrap();
let style_store = HwpxStyleStore::with_default_fonts("함초롬바탕");
let bytes = HwpxEncoder::encode(&validated, &style_store, &ImageStore::new()).unwrap();
std::fs::write("converted.hwpx", &bytes).unwrap();
}

CLI만 필요하다면 전용 명령도 이미 있습니다.

hwpforge convert-hwp5 legacy.hwp -o converted.hwpx
hwpforge audit-hwp5 legacy.hwp converted.hwpx
hwpforge census-hwp5 legacy.hwp --json

CLI에서 포맷 처리

CLI(hwpforge)는 변환·검사·편집·검증용 23개 명령을 제공합니다. 전체 목록은 hwpforge --help로 확인하세요. 이 장에서 다루는 포맷 변환 명령은 다음과 같습니다.

# Markdown → HWPX 변환
hwpforge convert report.md -o report.hwpx

# HWPX 문서 검사 (메타데이터 포함)
hwpforge inspect report.hwpx --json

# HWPX → JSON → 편집 → HWPX 라운드트립
hwpforge to-json report.hwpx -o report.json
# (AI 에이전트가 JSON 편집)
hwpforge from-json report.json -o updated.hwpx

# HWPX → Markdown (읽기용)
hwpforge to-md report.hwpx

# HWPX/HWP5 → PDF
hwpforge to-pdf report.hwpx -o report.pdf

HWPX → Markdown은 Rust API(MdEncoder::encode_lossy(&validated))로도 할 수 있습니다.

크레이트 역할 분담

크레이트역할포맷 의존성
hwpforge-foundation원시 타입 (HwpUnit, Color, Index)없음
hwpforge-core포맷 독립 문서 모델 (IR)없음
hwpforge-blueprintYAML 스타일 템플릿없음
hwpforge-smithy-hwpxHWPX ↔ Core 코덱HWPX (ZIP+XML)
hwpforge-smithy-hwp5HWP5 decode/projectionHWP5 (OLE/CFB)
hwpforge-smithy-mdMarkdown ↔ Core 코덱Markdown (텍스트)
hwpforge-smithy-pdfPDF 렌더러 (쓰기 전용, 코덱 아님)PDF (출력만)
hwpforge-convertHWP5 → HWPX 변환 오케스트레이터 + audit helpers + PDF 렌더 연산HWP5 + HWPX + PDF (smithy 경유)
hwpforge-bindings-cli / -mcp / -pyCLI / MCP / Python 진입점 (공유 연산 계층 hwpforge::ops 호출)직접 의존: CLI는 smithy 4종, MCP는 smithy-hwpx·smithy-md, Python은 smithy-pdf와 convert

핵심 원칙: Core 이하 계층은 어떤 파일 포맷도 모릅니다. 특정 포맷을 이해하는 것은 Smithy 계층이고(바인딩 크레이트는 그 Smithy에 직접 의존하기도 합니다), convert는 두 Smithy를 엮어 포맷 간 변환을 지휘합니다(자체 포맷 파싱 없음).

HWP5 포맷 구조 (참고)

HWP5 파일은 OLE2 Compound File Binary (CFB) 컨테이너 안에 바이너리 레코드 스트림을 저장합니다.

HWP5 파일 (OLE2 CFB)
├── FileHeader          — 파일 인식 정보, 버전, 플래그
├── DocInfo             — 문서 설정 (스타일, 폰트, 탭, 번호)
├── BodyText/
│   ├── Section0        — 첫 번째 섹션 (바이너리 레코드)
│   ├── Section1        — 두 번째 섹션
│   └── ...
├── BinData/            — 이미지 등 바이너리 데이터
├── DocOptions/         — 추가 옵션
├── Scripts/            — 매크로 스크립트
└── PrvText             — 미리보기 텍스트

각 섹션은 Tag-Length-Value (TLV) 구조의 레코드 체인으로 구성됩니다:

레코드 = TagID (10bit) + Level (10bit) + Size (12bit) + Data (Size bytes)

주의: HWP5의 TagID에는 +16 오프셋이 있습니다. PARA_HEADER = 0x42 (66), 공식 스펙의 0x32 (50)가 아닙니다.

현재 지원 상태

기능HWPXHWP5Markdown
읽기 (Decode)✅ 완전 지원🟡 전용 crate/CLI 경로✅ 완전 지원
쓰기 (Encode)✅ 완전 지원—✅ 완전 지원
메타데이터 추출✅ Core DOM🟡 inspect/census summary✅ YAML Frontmatter
이미지 추출✅ ImageStore🟡 decode/projection path—
스타일 보존✅ HwpxStyleStore🟡 warning-first projection + HWPX re-emission✅ StyleRegistry
JSON 라운드트립✅ to-json/from-json——

HWP5 읽기 자체는 더 이상 future tense가 아니다. 지금도 hwpforge-smithy-hwp5와 CLI를 통해 decode / inspect / convert / audit 경로를 사용할 수 있다. 다만 umbrella crate와 일부 top-level guide는 여전히 HWPX/Markdown 중심이며, HWP5 parity는 warning-first로 점진적으로 넓혀 가는 중이다.

대규모 HWP 아카이브 마이그레이션

레거시 HWP/HWPX 문서 아카이브를 검색 가능한 Markdown으로 마이그레이션하는 전략을 설명합니다.

마이그레이션 파이프라인 개요

┌────────────────┐     ┌──────────────┐     ┌─────────────┐     ┌──────────────┐
│ 1. 스캔        │ ──▶ │ 2. 분류      │ ──▶ │ 3. 변환     │ ──▶ │ 4. 검증      │
│ 파일 목록 수집 │     │ 포맷 감지    │     │ Core → MD   │     │ 무결성 확인  │
│                │     │ HWP5/HWPX    │     │ lossy 모드  │     │ 결과 기록    │
└────────────────┘     └──────────────┘     └─────────────┘     └──────────────┘

1단계: 파일 스캔 및 포맷 분류

파일 확장자와 매직 바이트로 포맷을 감지합니다.

이 장의 예제는 디렉터리 순회에 walkdir 크레이트를 씁니다. 따라 하려면 cargo add walkdir로 의존성을 추가하세요.

#![allow(unused)]
fn main() {
use std::path::{Path, PathBuf};

#[derive(Debug)]
enum DocFormat {
    Hwpx,         // ZIP + XML (PK 시그니처)
    Hwp5,         // OLE2/CFB (D0 CF 시그니처)
    Unknown(String),
}

#[derive(Debug)]
struct ScanResult {
    path: PathBuf,
    format: DocFormat,
    size_bytes: u64,
}

fn scan_archive(dir: &Path) -> Vec<ScanResult> {
    let mut results = Vec::new();

    let walker = walkdir::WalkDir::new(dir)
        .into_iter()
        .filter_map(|e| e.ok());

    for entry in walker {
        let path = entry.path();
        let ext = path.extension()
            .and_then(|e| e.to_str())
            .unwrap_or("")
            .to_lowercase();

        if ext != "hwp" && ext != "hwpx" {
            continue;
        }

        let size_bytes = entry.metadata().map(|m| m.len()).unwrap_or(0);
        let format = detect_format(path);

        results.push(ScanResult {
            path: path.to_path_buf(),
            format,
            size_bytes,
        });
    }

    results
}

fn detect_format(path: &Path) -> DocFormat {
    let Ok(bytes) = std::fs::read(path) else {
        return DocFormat::Unknown("읽기 실패".into());
    };

    if bytes.len() < 4 {
        return DocFormat::Unknown("파일이 너무 작음".into());
    }

    // ZIP (HWPX): PK\x03\x04
    if bytes.starts_with(&[0x50, 0x4B, 0x03, 0x04]) {
        return DocFormat::Hwpx;
    }

    // OLE2/CFB (HWP5): D0 CF 11 E0
    if bytes.starts_with(&[0xD0, 0xCF, 0x11, 0xE0]) {
        return DocFormat::Hwp5;
    }

    DocFormat::Unknown(format!("알 수 없는 시그니처: {:02X} {:02X}", bytes[0], bytes[1]))
}
}

2단계: 변환 (HWPX 중심 + HWP5 별도 경로)

기본 migration sample은 HWPX를 중심으로 설명합니다. 다만 현재는 HWP5도 전용 crate와 CLI로 decode, audit, HWPX re-emission 경로를 사용할 수 있습니다.

HWPX 파일 변환

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;
use hwpforge::md::MdEncoder;
use std::path::Path;

#[derive(Debug)]
struct ConvertResult {
    source: String,
    status: ConvertStatus,
    markdown_len: usize,
    title: Option<String>,
}

#[derive(Debug)]
enum ConvertStatus {
    Success,
    DecodeError(String),
    ValidationError(String),
    EncodeError(String),
}

fn convert_hwpx(input: &Path) -> ConvertResult {
    let source = input.display().to_string();

    // 1. 디코딩 (손상 파일 처리)
    let result = match HwpxDecoder::decode_file(input) {
        Ok(r) => r,
        Err(e) => {
            return ConvertResult {
                source,
                status: ConvertStatus::DecodeError(e.to_string()),
                markdown_len: 0,
                title: None,
            };
        }
    };

    let title = result.document.metadata().title.clone();

    // 2. 검증
    let validated = match result.document.validate() {
        Ok(v) => v,
        Err(e) => {
            return ConvertResult {
                source,
                status: ConvertStatus::ValidationError(e.to_string()),
                markdown_len: 0,
                title,
            };
        }
    };

    // 3. Markdown 변환 (RAG/검색용 lossy 모드)
    match MdEncoder::encode_lossy(&validated) {
        Ok(md) => ConvertResult {
            source,
            status: ConvertStatus::Success,
            markdown_len: md.len(),
            title,
        },
        Err(e) => ConvertResult {
            source,
            status: ConvertStatus::EncodeError(e.to_string()),
            markdown_len: 0,
            title,
        },
    }
}
}

HWP5 파일 처리

레거시 HWP5(.hwp) 파일도 현재 다룰 수 있습니다. 다만 대규모 migration에서는 HWPX 중심 파이프라인과 HWP5 전용 파이프라인을 분리하는 편이 운영이 쉽습니다.

전략설명자동화
HwpForge CLIconvert-hwp5, audit-hwp5, census-hwp5로 decode/점검/재출력완전 자동
전용 cratehwpforge-smithy-hwp5로 HWP5 decode 후 Core/HWPX 경로 재사용자동
별도 분류HWP5 파일만 분리해 별도 queue로 처리수동
// 현재 — HWP5 직접 decode 후 기존 pipeline에 연결
// use hwpforge_smithy_hwp5::Hwp5Decoder;
//
// let result = Hwp5Decoder::decode_file("legacy.hwp")?;
// let validated = result.document.validate()?;
// let markdown = MdEncoder::encode_lossy(&validated)?;

3단계: 배치 처리 아키텍처

대규모 아카이브(수천~수만 파일)를 안정적으로 처리하는 패턴입니다.

#![allow(unused)]
fn main() {
use std::path::{Path, PathBuf};
use std::fs;

use hwpforge::hwpx::HwpxDecoder;
use hwpforge::md::MdEncoder;

struct MigrationConfig {
    input_dir: PathBuf,
    output_dir: PathBuf,
    error_dir: PathBuf,
    /// 개별 파일 처리 제한 시간 (초)
    timeout_secs: u64,
    /// 최대 파일 크기 (바이트, 기본 100MB)
    max_file_size: u64,
}

struct MigrationReport {
    total: usize,
    success: usize,
    failed: usize,
    skipped_hwp5: usize,
    skipped_too_large: usize,
    errors: Vec<(String, String)>,
}

fn run_migration(config: &MigrationConfig) -> MigrationReport {
    fs::create_dir_all(&config.output_dir).expect("출력 디렉토리 생성 실패");
    fs::create_dir_all(&config.error_dir).expect("오류 디렉토리 생성 실패");

    let mut report = MigrationReport {
        total: 0, success: 0, failed: 0,
        skipped_hwp5: 0, skipped_too_large: 0,
        errors: Vec::new(),
    };

    let files: Vec<_> = walkdir::WalkDir::new(&config.input_dir)
        .into_iter()
        .filter_map(|e| e.ok())
        .filter(|e| {
            e.path().extension()
                .is_some_and(|ext| ext == "hwpx" || ext == "hwp")
        })
        .collect();

    report.total = files.len();
    eprintln!("총 {} 파일 발견", report.total);

    for (i, entry) in files.iter().enumerate() {
        let path = entry.path();
        let rel_path = path.strip_prefix(&config.input_dir).unwrap_or(path);

        // 진행률 표시
        if (i + 1) % 100 == 0 || i + 1 == report.total {
            eprintln!("[{}/{}] 처리 중...", i + 1, report.total);
        }

        // 예시 단순화를 위해 HWP5는 별도 queue로 분리
        if path.extension().is_some_and(|ext| ext == "hwp") {
            report.skipped_hwp5 += 1;
            continue;
        }

        // 파일 크기 제한
        let size = entry.metadata().map(|m| m.len()).unwrap_or(0);
        if size > config.max_file_size {
            report.skipped_too_large += 1;
            continue;
        }

        // 변환 시도
        match convert_single(path, &config.output_dir, rel_path) {
            Ok(_) => report.success += 1,
            Err(e) => {
                report.failed += 1;
                report.errors.push((path.display().to_string(), e.clone()));

                // 실패 파일을 오류 디렉토리에 복사
                let err_dest = config.error_dir.join(rel_path);
                if let Some(parent) = err_dest.parent() {
                    let _ = fs::create_dir_all(parent);
                }
                let _ = fs::copy(path, err_dest);
            }
        }
    }

    report
}

fn convert_single(
    input: &Path,
    output_dir: &Path,
    rel_path: &Path,
) -> Result<(), String> {
    let result = HwpxDecoder::decode_file(input)
        .map_err(|e| format!("디코딩 실패: {e}"))?;

    let validated = result.document.validate()
        .map_err(|e| format!("검증 실패: {e}"))?;

    let markdown = MdEncoder::encode_lossy(&validated)
        .map_err(|e| format!("MD 인코딩 실패: {e}"))?;

    // 출력 경로: .hwpx → .md
    let out_name = rel_path.with_extension("md");
    let out_path = output_dir.join(out_name);

    if let Some(parent) = out_path.parent() {
        fs::create_dir_all(parent).map_err(|e| format!("디렉토리 생성 실패: {e}"))?;
    }

    fs::write(&out_path, &markdown).map_err(|e| format!("파일 쓰기 실패: {e}"))?;

    Ok(())
}
}

4단계: 무결성 검증

변환 결과의 품질을 검증합니다.

기본 검증 항목

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;
use hwpforge::md::MdEncoder;
use std::path::Path;

struct IntegrityCheck {
    source_sections: usize,
    source_paragraphs: usize,
    source_tables: usize,
    markdown_lines: usize,
    markdown_bytes: usize,
    has_title: bool,
}

fn verify_conversion(hwpx_path: &Path, markdown: &str) -> Option<IntegrityCheck> {
    let result = HwpxDecoder::decode_file(hwpx_path).ok()?;
    let doc = &result.document;

    let mut total_paragraphs = 0;
    let mut total_tables = 0;
    for section in doc.sections() {
        total_paragraphs += section.paragraphs.len();
        total_tables += section.content_counts().tables;
    }

    Some(IntegrityCheck {
        source_sections: doc.sections().len(),
        source_paragraphs: total_paragraphs,
        source_tables: total_tables,
        markdown_lines: markdown.lines().count(),
        markdown_bytes: markdown.len(),
        has_title: doc.metadata().title.is_some(),
    })
}
}

검증 체크리스트

항목방법허용 기준
텍스트 보존원본 문단 수 vs Markdown 비어있지 않은 줄 수손실 < 10%
표 구조원본 표 수 vs Markdown | 테이블 수동일
메타데이터YAML frontmatter에 title/author 존재원본과 일치
파일 크기Markdown 바이트 > 0빈 파일 없음
인코딩UTF-8 유효성깨진 문자 없음

손상 파일 처리 전략

대규모 아카이브에서는 손상되거나 비표준 파일이 불가피합니다.

일반적인 오류 유형과 대응

오류원인대응
ZIP 파싱 실패파일 손상, 불완전 다운로드오류 목록에 기록, 원본 보존
XML 파싱 실패비표준 네임스페이스, 잘못된 인코딩오류 목록에 기록, 수동 검토
검증 실패빈 섹션, 유효하지 않은 인덱스경고 후 계속 진행
OOM (메모리 부족)매우 큰 임베디드 이미지파일 크기 제한으로 사전 필터링
암호화된 파일비밀번호 보호별도 목록으로 분류

에러 리포트 생성

#![allow(unused)]
fn main() {
use std::fs;

struct MigrationReport {
    total: usize,
    success: usize,
    failed: usize,
    skipped_hwp5: usize,
    skipped_too_large: usize,
    errors: Vec<(String, String)>,
}

fn write_report(report: &MigrationReport, path: &str) {
    let mut lines = Vec::new();
    lines.push(format!("# 마이그레이션 리포트\n"));
    lines.push(format!("- 총 파일: {}", report.total));
    lines.push(format!("- 성공: {}", report.success));
    lines.push(format!("- 실패: {}", report.failed));
    lines.push(format!("- HWP5 별도 처리: {}", report.skipped_hwp5));
    lines.push(format!("- 크기 초과: {}", report.skipped_too_large));

    if !report.errors.is_empty() {
        lines.push(format!("\n## 실패 목록\n"));
        lines.push(format!("| 파일 | 오류 |"));
        lines.push(format!("| --- | --- |"));
        for (file, err) in &report.errors {
            lines.push(format!("| `{}` | {} |", file, err));
        }
    }

    fs::write(path, lines.join("\n")).expect("리포트 저장 실패");
}
}

Lossy vs Lossless 모드 선택

목적권장 모드이유
RAG/검색 인덱싱encode_lossy표준 GFM, 청크 분할 호환, 토큰 절약
아카이브 백업encode_lossless구조 완전 보존, 원본 복원 가능
하이브리드둘 다 생성lossy는 검색용, lossless는 백업용

하이브리드 전략이 이상적입니다:

#![allow(unused)]
fn main() {
use hwpforge::hwpx::HwpxDecoder;
use hwpforge::md::MdEncoder;

let result = HwpxDecoder::decode_file("document.hwpx").unwrap();
let validated = result.document.validate().unwrap();

// 검색/RAG용
let lossy = MdEncoder::encode_lossy(&validated).unwrap();
std::fs::write("output/search/document.md", &lossy).unwrap();

// 아카이브 백업용
let lossless = MdEncoder::encode_lossless(&validated).unwrap();
std::fs::write("output/archive/document.lossless.md", &lossless).unwrap();
}

관련 문서

CLI 레퍼런스

hwpforge(Hammer)는 HwpForge의 명령줄 도구입니다. 이 장은 hwpforge --help와 hwpforge <명령> --help가 보여 주는 23개 명령을 정리합니다. 플래그 이름과 기본값은 clap 정의(crates/hwpforge-bindings-cli/src/main.rs)를 그대로 옮겼으므로, 버전이 올라 의심스러우면 --help가 우선입니다.

설치

hwpforge-bindings-cli는 crates.io에 배포하지 않으므로(publish = false) git 또는 로컬 경로에서 설치합니다. hwpforge-smithy-pdf(krilla) 의존 때문에 워크스페이스 MSRV(1.89)보다 높은 Rust 1.92 이상이 필요합니다(rust-version = "1.92").

cargo install --git https://github.com/ai-screams/HwpForge hwpforge-bindings-cli

# 또는 clone 후 로컬 경로에서
git clone https://github.com/ai-screams/HwpForge && cd HwpForge
cargo install --path crates/hwpforge-bindings-cli

설치되는 바이너리 이름은 hwpforge입니다([[bin]] name = "hwpforge").

전역 옵션

옵션설명
--json결과를 machine-readable JSON으로 출력합니다
-h, --help도움말
-V, --version버전

--json은 모든 명령이 받습니다. 에이전트가 결과를 파싱할 때는 이 플래그를 붙이세요.

명령 한눈에 보기

--help는 HWP5·PDF 명령을 먼저 나열하고 그 뒤에 HWPX 명령을 나열합니다. 아래 표는 용도별로 다시 묶었습니다.

용도명령
HWP5 · PDFaudit-hwp5, census-hwp5, convert-hwp5, to-pdf
만들기 · 변환convert, to-json, from-json, to-md
읽기 · 검사inspect, outline, fields, read, diff, validate
편집fill, set-cell, patch, insert-para, delete-para, stamp-plan, stamp
부가templates, schema

HWP5 · PDF

convert-hwp5

HWP5(.hwp)를 HWPX로 변환합니다.

hwpforge convert-hwp5 [OPTIONS] --output <OUTPUT> <INPUT>
옵션설명
-o, --output출력 HWPX 경로(필수)
--carry-layout-cacheHWP5의 조판 캐시(PARA_LINE_SEG)를 HWPX <hp:linesegarray>로 실어 to-pdf로 렌더할 수 있게 합니다

--carry-layout-cache로 만든 결과는 PDF 재생·대조 전용입니다. 한컴에서 다시 여는 용도로 쓰면 여러 줄이 겹칠 위험이 있다고 --help가 밝힙니다. 좌표를 정규화할 수 없는 문단(차트, 알 수 없는 컨트롤, 모호한 마커 경계)은 잘못된 데이터를 싣는 대신 LAYOUT_CACHE_DROPPED 경고와 함께 캐시를 버립니다.

to-pdf

HWPX 또는 HWP5 문서를 조판 캐시 재생 방식으로 PDF로 렌더합니다. 입력 형식은 확장자가 아니라 내용으로 판별합니다.

hwpforge to-pdf [OPTIONS] <INPUT>
옵션설명
-o, --output출력 PDF 경로(기본값: 입력 경로의 확장자를 .pdf로 바꾼 것)
--font-dir <FONT_DIRS>폰트를 찾을 디렉터리(반복 가능)
--discovery <DISCOVERY>폰트 자동 탐색: explicit(결정적, 기본값) | hancom | platform
--degraded렌더 실패를 오류 대신 경고로 낮춥니다. 없는 스타일·축 폰트는 regular로 그리고, 렌더할 수 없는 이미지는 건너뜁니다
--partial-cache-reject조판 캐시가 없는 문단이 하나라도 있으면 문서를 거부합니다(기본값: 경고 후 그 문단을 건너뜀)

조판 캐시가 있는 문서만 렌더할 수 있습니다. convert나 from-json이 새로 만든 문서에는 캐시가 없어 거부됩니다.

audit-hwp5

HWP5 원본과 변환된 HWPX 결과의 구조·의미 동등성을 점검합니다. 픽셀 단위 시각·쪽 나눔 동등성은 보장하지 않으며, 후속 시각 검증용 체크리스트를 보고서에 담습니다.

hwpforge audit-hwp5 [OPTIONS] <SOURCE> <RESULT>

<SOURCE>는 원본 .hwp, <RESULT>는 변환된 .hwpx입니다.

census-hwp5

HWP5 파일(과 선택적 HWPX 짝 파일)의 원시 fixture census를 만듭니다.

hwpforge census-hwp5 [OPTIONS] <INPUT>
옵션설명
--companion <COMPANION>중첩 XML·경로 census용 HWPX 짝 fixture(선택)
-o, --outputcensus 데이터셋을 JSON으로 저장할 경로(선택)

만들기 · 변환

convert

Markdown을 HWPX로 변환합니다.

hwpforge convert [OPTIONS] --output <OUTPUT> <INPUT>
옵션설명
<INPUT>Markdown 파일. -를 주면 stdin에서 읽습니다
-o, --output출력 HWPX 경로(필수)
--preset <PRESET>스타일 프리셋 이름(기본값: default)

사용 가능한 프리셋은 templates list로 확인합니다.

to-json

HWPX를 편집 가능한 JSON으로 내보냅니다.

hwpforge to-json [OPTIONS] --output <OUTPUT> <FILE>
옵션설명
-o, --output출력 JSON 경로(필수, stdout 내보내기는 없습니다)
--section <SECTION>특정 섹션만 추출(0부터 시작하는 인덱스)
--no-styles스타일 정보를 제외합니다

--section 없이 내보낸 문서 전체 JSON은 from-json이, --section N으로 내보낸 섹션 JSON은 patch가 읽습니다. 두 JSON은 서로 다른 형식입니다.

from-json

JSON을 HWPX로 되돌립니다.

hwpforge from-json [OPTIONS] --output <OUTPUT> <INPUT>
옵션설명
-o, --output출력 HWPX 경로(필수)
--base <BASE>이미지를 물려받을 기준 HWPX(왕복 충실도용)

to-md

HWPX를 Markdown으로 변환합니다.

hwpforge to-md [OPTIONS] <INPUT>
옵션설명
-o, --output출력 디렉터리(기본값: 입력과 같은 디렉터리)
--mode <MODE>styled(기본값, 스타일을 따르고 이미지 포함) | lossy(스타일 정보 없는 읽기용) | lossless(YAML 프론트매터가 붙은 왕복 안전 형식)

읽기 · 검사

inspect

디코드한 HWPX 문서 구조를 보여 줍니다. HWPX만 받습니다. HWP5는 먼저 convert-hwp5로 변환하거나 audit-hwp5로 비교하세요.

hwpforge inspect [OPTIONS] <FILE>
옵션설명
--styles스타일 상세(문자·문단 모양)를 포함합니다

outline

문서의 내비게이션 맵(제목, 표, 이름 있는 누름틀, 책갈피)을 보여 줍니다. 이름 앵커(제목 텍스트, 표 순번, 필드·책갈피 이름)가 1차 키이고 {section, para} 위치는 구조 편집 뒤에 낡을 수 있으므로, 읽기·편집 전에 한 번 받아 두는 용도입니다.

hwpforge outline [OPTIONS] <FILE>

fields

이름 있는 누름틀(click-here field)과 각각이 채울 수 있는 필드인지 나열합니다.

hwpforge fields [OPTIONS] <FILE>

read

문서 전체를 내보내지 않고 목표 하나만 텍스트로 읽습니다. 목표는 정확히 하나만 지정합니다.

hwpforge read [OPTIONS] <FILE>
옵션설명
--section <SECTION>문단을 읽을 섹션 인덱스
--paras <PARAS>닫힌 문단 범위 "A..B" 또는 단일 "N"(--section 필요)
--table <TABLE>표 순번. 논리 격자 텍스트 행렬로 읽습니다(병합 영역은 앵커에 한 번, span과 함께)
--field <FIELD>이름 있는 누름틀 하나

텍스트가 아닌 내용(표, 이미지, 컨트롤)은 조용히 버려지지 않고 명시적인 표지로 드러납니다. 읽기 전용입니다.

diff

두 HWPX를 두 채널로 비교해 편집이 실제로 무엇을 바꿨는지 확인합니다.

hwpforge diff [OPTIONS] <BASE> <REVISED>
옵션설명
-o, --output전체 JSON 보고서를 이 경로에도 씁니다

semantic 채널은 디코드한 Core 구조를 필드 값, 표 셀 텍스트 {table, row, col}, 문단 텍스트 {section, para}, 구조 개수, 상한이 있는 미분류 잔여로 분류합니다. package 채널은 ZIP 엔트리를 바이트로 비교합니다. 바뀐 엔트리 안의 wire 내용(예: hp:linesegarray 조판 캐시)은 항목별로 나열하지 않으며 보고서가 그렇게 밝힙니다. fill·set-cell·stamp·patch 뒤에 돌려 의도한 변경만 반영됐는지 확인하세요.

validate

HWPX가 디코드되고 Core의 구조 불변식(Document::validate)을 만족하는지 편집 없이 검사합니다.

hwpforge validate [OPTIONS] <FILE>

종료 코드는 --help에 다음과 같이 정의돼 있습니다.

종료 코드의미
0문서가 유효합니다
1파일을 읽을 수 없음(없거나 읽을 수 없는 파일)
2바이트가 디코드 가능한 HWPX 패키지가 아닙니다
3디코드는 되지만 검증에 실패합니다(오류가 아니라 판정이며, --json 출력의 errors를 보세요)

인자를 잘못 준 경우(필수 인자 누락, 알 수 없는 옵션)는 clap이 종료 코드 2로 끝냅니다. 도움말은 이를 1로 적고 있으나 실측은 2입니다(validate, validate --bogus x 모두 2). 이 문단은 validate에 대한 것이며 다른 명령의 종료 코드는 여기서 다루지 않습니다.

편집

fill·set-cell·insert-para·delete-para·stamp·patch는 결과를 -o로 지정한 출력 경로에 씁니다(필수 옵션). stamp-plan은 파일을 쓰지 않으므로 -o가 없습니다. 편집 뒤에는 diff로 변경 범위를 확인하세요.

fill

이름 있는 누름틀을 값으로 채웁니다. 건드리지 않는 패키지 엔트리는 바이트 그대로 보존합니다.

hwpforge fill [OPTIONS] --output <OUTPUT> <FILE>
옵션설명
--set <NAME=VALUE>채울 이름=값 쌍(반복 가능). 하나 이상 필수
-o, --output출력 HWPX 경로(필수)

요청한 값을 모두 먼저 검증(알 수 없는·중복된·채울 수 없는 이름, 빈 값)하고, 하나라도 실패하면 아무것도 쓰지 않습니다. 누름틀이 들어 있는 서식 문서에서만 동작하므로 convert로 막 만든 문서에는 필드가 없습니다. 이름은 fields로 먼저 확인하세요.

set-cell

표 셀을 논리 격자 주소로 편집합니다. 실패 시 아무것도 쓰지 않는 admission 게이트 뒤에서 동작합니다.

hwpforge set-cell [OPTIONS] --output <OUTPUT> <FILE>
옵션설명
--table <TABLE>표 순번(문서 순서, 0부터). 단일 편집에는 필수(--map 일괄 편집에서는 쓰지 않음)
--at <R,C>격자 좌표 "row,col"(덮인 위치는 병합 앵커로 해석)
--right-of <LABEL>유일하게 라벨된 셀의 오른쪽 셀을 대상으로 합니다
--below <LABEL>유일하게 라벨된 셀의 아래쪽 셀을 대상으로 합니다
--text <TEXT>바꿀 텍스트(빈 문자열은 셀을 비웁니다). 단일 편집에는 필수
--map <MAP_JSON>일괄 편집용 명세 맵(CellSpec JSON 배열). 위 플래그들과 함께 쓸 수 없습니다
-o, --output출력 HWPX 경로(필수)

patch

기존 HWPX의 섹션 하나를 교체합니다.

hwpforge patch [OPTIONS] --section <SECTION> --output <OUTPUT> <BASE> <SECTION_JSON>

<SECTION_JSON>은 to-json --section N으로 내보낸 섹션 JSON이며, <BASE>가 이미지와 나머지 패키지의 기준입니다. 문서 전체 JSON을 넣으면 스키마 불일치로 거부됩니다. 구조(문단 수)를 바꾸려면 insert-para·delete-para나 from-json을 쓰세요.

insert-para

기준 문단 앞이나 뒤에 새 최상위 문단을 삽입합니다. 다른 바이트는 그대로 둡니다.

hwpforge insert-para [OPTIONS] --section <SECTION> --anchor <ANCHOR> --text <TEXTS> --output <OUTPUT> <FILE>
옵션설명
--section <SECTION>섹션 인덱스
--anchor <ANCHOR>기준 문단 인덱스(새 문단이 이 문단의 문단·문자 모양을 그대로 물려받습니다)
--before기준 앞에 삽입합니다(기본값은 뒤)
--text <TEXTS>새 문단의 한 줄 일반 텍스트. 반복하면 연속된 블록을 한 번의 검증된 편집으로 삽입합니다
-o, --output출력 HWPX 경로(필수)

섹션의 첫 문단(secPr를 담은 문단) 앞 삽입은 거부됩니다. 왕복 안전한 입력만 편집할 수 있습니다.

delete-para

최상위 본문 문단을 인덱스로 삭제합니다. 전부 성공하거나 전부 취소됩니다.

hwpforge delete-para [OPTIONS] --section <SECTION> --output <OUTPUT> <FILE>
옵션설명
--section <SECTION>섹션 인덱스
--index <INDICES>...삭제할 최상위 문단 인덱스(일괄 삭제는 반복). 하나 이상 필수
-o, --output출력 HWPX 경로(필수)

참조(책갈피·상호참조·각주 등)를 담은 문단, 강제 쪽·단 나눔이 있는 문단, 섹션 속성(secPr)을 담은 첫 문단, 삭제하면 섹션이 비게 되는 경우는 거부합니다(fail-closed).

stamp-plan

산문 속 자리표시자(체크박스, 괄호 빈칸, 날짜 빈칸, 단독 @, 도장 토큰) 후보를 찾아 템플릿 스탬핑 계획을 만듭니다.

hwpforge stamp-plan [OPTIONS] <FILE>

출력의 각 후보에 {"field":{"name":"…"}} 또는 "ignore" 액션을 붙여 명세 맵을 만든 뒤 stamp --map을 실행합니다. 지침 문맥의 후보(guarded)는 자동 적용되지 않습니다.

stamp

명세 맵을 적용해 자리표시자를 이름 있는 누름틀로 승격합니다. 전부 성공하거나 전부 취소되며, 스탬프된 HWPX와 매니페스트를 씁니다.

hwpforge stamp [OPTIONS] --map <MAP_JSON> --output <OUTPUT> <FILE>
옵션설명
--map <MAP_JSON>StampSpec JSON 배열(필수)
-o, --output출력 HWPX 경로(필수)
--manifest <MANIFEST>매니페스트 JSON 경로(기본값: <output>.manifest.json)

실패 시 아무것도 쓰지 않는 admission 게이트(무변경 왕복 + ZIP 닫힌 세계 검사) 뒤에서 동작하고, 결과물은 곧바로 fields·fill에 쓸 수 있습니다.

부가

templates

스타일 프리셋을 관리합니다. 하위 명령은 두 개입니다.

hwpforge templates list            # 사용 가능한 프리셋 목록
hwpforge templates show <NAME>     # 프리셋 상세

schema

문서·스타일 타입의 JSON Schema를 출력합니다.

hwpforge schema [OPTIONS] [TYPE_NAME]

TYPE_NAME은 document(기본값), exported-document, exported-section 중 하나입니다. exported-document가 to-json의 전체 JSON, exported-section이 --section JSON의 형식입니다.

다음 단계

MCP 서버 레퍼런스

hwpforge-mcp는 HwpForge의 MCP 서버입니다. Claude Code 같은 MCP 지원 AI 도구가 한글 문서를 직접 만들고 읽고 편집할 수 있도록 도구 19개, 리소스 4개, 프롬프트 3개를 노출합니다. 이 장의 이름과 매개변수는 crates/hwpforge-bindings-mcp/src의 정의에서 옮겼습니다.

모든 도구는 { data, summary, next } 3층 출력 형식을 씁니다. 파일은 대부분 경로로 주고받으며(예외: hwpforge_convert는 is_file: false로 Markdown을 인라인으로 받고, hwpforge_to_json은 output_path 없이 JSON을 인라인으로 돌려주고, hwpforge_diff는 output_path 없이 보고서를 인라인으로 돌려주며, hwpforge_from_json은 JSON을 structure 문자열로 받습니다), 실패는 오류 응답(CallToolResult::error)으로 돌아옵니다.

등록

npm (권장, Rust 툴체인 불필요)

npx -y가 플랫폼에 맞는 바이너리를 내려받습니다.

# 현재 프로젝트에서만
claude mcp add hwpforge -- npx -y @hwpforge/mcp

# 모든 프로젝트에서 (user 범위)
claude mcp add --scope user hwpforge -- npx -y @hwpforge/mcp

프로젝트 루트의 .mcp.json에 직접 적어도 됩니다.

{
  "mcpServers": {
    "hwpforge": {
      "command": "npx",
      "args": ["-y", "@hwpforge/mcp"]
    }
  }
}

Claude Code의 MCP 서버는 claude mcp add 또는 .mcp.json으로 등록합니다. .claude/settings.json에 적어도 MCP 서버로 읽히지 않습니다.

Cargo

cargo install hwpforge-bindings-mcp
claude mcp add hwpforge hwpforge-mcp

crates.io 패키지 이름은 hwpforge-bindings-mcp이고 설치되는 바이너리 이름은 hwpforge-mcp입니다([[bin]] name = "hwpforge-mcp").

npm 패키지 구성

@hwpforge/mcp는 플랫폼별 패키지를 optionalDependencies로 가진 기본 패키지입니다. 플랫폼 패키지는 .github/workflows/npm-publish.yml의 빌드 매트릭스와 같은 다섯 개입니다.

플랫폼 패키지 접미사Rust 타깃
darwin-arm64aarch64-apple-darwin
darwin-x64x86_64-apple-darwin
linux-x64x86_64-unknown-linux-gnu
linux-arm64aarch64-unknown-linux-gnu
win32-x64x86_64-pc-windows-msvc

패키지 이름은 @hwpforge/mcp-<접미사> 형태입니다. 게시는 npm Trusted Publishing(OIDC)만 씁니다.

도구

표의 (.hwpx)는 그 output_path가 .hwpx로 끝나야 한다는 표시입니다. 이 확장자 검사는 hwpforge_convert·hwpforge_from_json·hwpforge_restyle·hwpforge_fill·hwpforge_set_cell·hwpforge_patch·hwpforge_stamp에 있고, hwpforge_insert_para·hwpforge_delete_para에는 없습니다. hwpforge_to_json의 output_path는 .json으로 끝나야 합니다.

만들기 · 변환

도구용도매개변수
hwpforge_convertMarkdown을 HWPX로 변환합니다markdown(파일 경로 또는 인라인 내용), is_file(기본 true), output_path(.hwpx), preset(기본 default)
hwpforge_to_jsonHWPX를 편집용 JSON으로 내보냅니다file_path, section(선택, 0부터), output_path(선택, .json으로 끝나야 함, 없으면 JSON을 응답에 인라인으로 돌려줌)
hwpforge_from_jsonJSON(ExportedDocument 스키마)으로 HWPX를 직접 만듭니다structure(JSON 문자열), output_path(.hwpx)
hwpforge_to_mdHWPX를 Markdown으로 변환합니다file_path, output_dir(선택, 기본값은 입력과 같은 디렉터리)
hwpforge_restyle기존 HWPX에 다른 스타일 프리셋을 적용합니다file_path, preset, output_path(.hwpx)
hwpforge_templates스타일 프리셋 목록을 돌려줍니다name(선택, 프리셋 이름 필터)

hwpforge_to_json의 인라인 응답은 직렬화된 응답이 1 MB 미만일 때만 가능하며, 더 큰 내보내기는 OUTPUT_TOO_LARGE로 거부되므로 output_path를 주세요. 전체 문서 내보내기는 hwpforge_from_json이, section을 준 내보내기는 hwpforge_patch가 읽습니다.

읽기 · 검사

도구용도매개변수
hwpforge_inspect섹션·문단·표·이미지·차트·머리글/바닥글·쪽번호 개수와 디코드 경고를 돌려줍니다file_path, styles(예약 필드, 현재 무시됨)
hwpforge_outline제목, 표, 이름 있는 누름틀, 책갈피의 내비게이션 맵file_path
hwpforge_fields누름틀의 이름, 힌트, 현재 값, 채울 수 있는지 여부file_path
hwpforge_read문단 범위, 표 격자, 필드 중 하나만 텍스트로 읽습니다file_path, section, paras("A..B" 또는 "N"), table, field — section/table/field 중 정확히 하나
hwpforge_diff두 HWPX를 semantic·package 두 채널로 비교합니다base_path, revised_path, output_path(선택, 보고서가 인라인 1 MB를 넘으면 필수)
hwpforge_validateHWPX 구조와 무결성을 검사합니다file_path

hwpforge_validate는 디코드할 수 없는 파일(예: .hwp)을 “유효하지 않은 문서“가 아니라 DECODE_FAILED 오류로 알립니다.

편집

도구용도매개변수
hwpforge_fill이름 있는 누름틀을 이름→값 맵으로 채웁니다(전부 성공하거나 전부 취소)file_path, values(이름→값 맵), output_path(.hwpx)
hwpforge_set_cell표 셀을 논리 격자 주소로 편집합니다file_path, specs(셀 명세 배열: 표 순번 + at {row,col} / right_of / below + text), output_path(.hwpx)
hwpforge_patch섹션 하나를 편집한 JSON으로 교체합니다(텍스트 전용)base_path, section, section_json_path, output_path(.hwpx)
hwpforge_insert_para기준 문단 앞뒤에 새 최상위 문단을 삽입합니다file_path, section, anchor, before(기본 false), text 또는 texts 중 정확히 하나, output_path
hwpforge_delete_para최상위 본문 문단을 인덱스로 삭제합니다file_path, section, indices, output_path
hwpforge_stamp_plan자리표시자 후보를 찾습니다file_path
hwpforge_stamp승인된 명세로 자리표시자를 누름틀로 승격합니다file_path, specs(텍스트 명세), cells(셀 명세), source_sha256(cells를 쓰면 필수), output_path(.hwpx), manifest_path(기본 <output>.manifest.json)

편집 도구의 동작 규칙은 CLI 레퍼런스의 대응하는 명령(예: hwpforge_insert_para는 insert-para)과 같습니다. 주의할 점은 다음과 같습니다.

  • hwpforge_patch는 문단 구조를 바꾸지 못합니다. 의미 텍스트 슬롯의 개수나 경로가 다르면 거부되며, 문단 추가·삭제는 hwpforge_insert_para·hwpforge_delete_para, 표 셀은 hwpforge_set_cell, 구조가 바뀐 문서 재구성은 hwpforge_from_json을 씁니다.
  • hwpforge_delete_para·hwpforge_insert_para·hwpforge_set_cell·hwpforge_stamp는 왕복 안전한 입력만 편집합니다. 인코더가 ZIP 엔트리를 모두 실어 보낼 수 없는 문서, 곧 한컴이 저장하며 Preview/*와 META-INF/container.rdf를 더한 문서는 거부됩니다. 이런 문서에도 hwpforge_to_json + hwpforge_patch(텍스트)와 hwpforge_fill(누름틀)은 쓸 수 있습니다.
  • hwpforge_stamp는 hwpforge_stamp_plan이 준 후보 객체를 그대로 복사해 action({"field":{"name":"…"}} 또는 "ignore")만 더합니다. 지침 문맥의 후보는 자동 적용되지 않습니다.

리소스

스타일 템플릿을 YAML(application/x-yaml)로 읽는 리소스 4개입니다. 프리셋 이름과 일치합니다(crates/hwpforge-bindings-mcp/src/resources/mod.rs).

URI이름
hwpforge://templates/defaultDefault Template
hwpforge://templates/modernModern Template
hwpforge://templates/classicClassic Template
hwpforge://templates/latestLatest Template

프롬프트

워크플로 안내 프롬프트 3개입니다(crates/hwpforge-bindings-mcp/src/prompts/mod.rs).

이름제목인자
generate_proposal정부 제안서 생성topic(필수), organization(선택), deadline(선택, YYYY-MM-DD)
generate_report보고서 생성topic(필수), author(선택), report_type(선택: research / progress / analysis, 기본 research)
convert_and_review문서 편집 워크플로우file_path(필수), edit_instructions(선택)

참고

  • MCP 서버는 베타이며 HWP5 경로는 MCP가 아니라 CLI(convert-hwp5, audit-hwp5, to-pdf)를 우선합니다. 이 서버의 19개 도구에는 HWP5와 PDF 관련 도구가 없습니다.
  • CLI 명령은 CLI 레퍼런스, Python 사용법은 Python 가이드를 보세요.

API 레퍼런스 (rustdoc)

HwpForge의 전체 공개 API 문서는 docs.rs에서 확인할 수 있습니다.

참고

  • 이 mdBook는 개념 설명과 사용 가이드 중심입니다.
  • trait, struct, enum, function의 상세 시그니처는 rustdoc이 진실입니다.
  • docs.rs와 저장소 문서가 어긋나면, 공개 API 계약은 rustdoc 쪽을 먼저 확인하십시오.

Changelog

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.

릴리스 노트 위치 (0.10.0 이후)

0.10.0 이후 릴리스 노트는 release-plz 가 크레이트별로 생성한다 — 이 루트 파일은 갱신을 멈췄다. 최신 이력은 다음에서 확인:

아래는 0.6.0~0.9.0 수동 관리 이전 이력이다.

[0.6.0 – 0.9.0] — 2026-05-29 … 2026-06-28 (released)

아래 항목들은 이미 crates.io 로 릴리스된 v0.6.0~v0.9.0 누적분이다 (루트 CHANGELOG 를 버전별 섹션으로 컷하지 못해 오랫동안 [Unreleased] 에 쌓여 있던 것 — E6 IR 와이어-누출 상환 A/B/C/M2, Wave 11/12 HWP5→HWPX carry, 문서 메타데이터 등). 정확한 버전별 대응은 release-plz 가 소유하는 per-crate CHANGELOG (crates/*/CHANGELOG.md) 를 참조.

Changed — BREAKING — cross-ref inst_id 누출을 ObjectId 로 (E6/M2, ADR-010)

크로스레퍼런스 링크가 두 무관한 정수 필드(타깃 inst_id: Option<u64/u32> ↔ 참조자 RefTarget::SystemId(u64))의 값-우연-일치였던 것을, 공유 newtype hwpforge_core::ObjectId(u64) 로 묶어 타입 수준 링크로 격상 (ADR-010, ADR-005 supersede). 리팩토링이 링크를 조용히 깨면 byte-diff 로도 안 잡히던 취약성 제거 + notes(u32)/shapes(u64) 폭 불일치 해소.

  • 신규 public 타입 hwpforge_core::ObjectId (#[serde(transparent)] → JSON/YAML 에서 bare integer).
  • 타깃 inst_id 필드 타입 변경 (필드명 유지): Image/Table/ Control::{Equation,Group,TextArt,Footnote,Endnote} 의 inst_id → Option<ObjectId>. Footnote/Endnote 는 u32→u64 승격. 패턴 매칭에서 값을 꺼내 쓰면 ObjectId 로 받게 됨(Some(42) → Some(ObjectId::new(42))).
  • RefTarget::SystemId(u64) → RefTarget::Object(ObjectId) (variant 리네임
    • payload 타입). FromStr/Display 출력 wire 문자열(#<id>)은 불변.
  • Control::{footnote,endnote}_with_id 인자 u32 → u64 (정수 리터럴 호출은 무변경).
  • JSON 영향: 타깃 inst_id 값은 정수 그대로(byte-동일). RefTarget 직렬화 태그가 "SystemId" → "Object" 로 변경 (CLI to-json/from-json, MCP). wire(HWPX) 출력은 전(全) 구간 byte-중립.
  • 와이어 스키마 HxFootNote.inst_id u32→u64 (in-range byte-중립, truncation 제거).

Fixed — Markdown(GFM) 표 헤더 행 손실 (사용자 흐름 점검)

convert(Markdown → HWPX) 에서 GFM 표의 헤더 행이 통째로 사라지던 data-loss 버그. | A | B | 헤더 + | 1 | 2 | 데이터 → HWPX 에 1|2 1행만 남고 A|B 소실.

  • 원인: md 디코더가 pulldown-cmark 의 Tag::TableHead/TagEnd::TableHead 를 no-op({}) 처리 → 헤더 셀이 어느 행에도 안 붙고 드롭. 본문 행(TableRow)만 캡처.
  • 수정: TableHead 를 TableRow 처럼 처리해 헤더 셀을 표 행 0 으로 캡처 (Core/HWPX 는 행 0 을 헤더로 렌더 — to-md 가 행 0 을 md 헤더로 출력하는 것과 정합).
  • 검증: convert → HWPX 에 2행(A,B / 1,2) + to-md 라운드트립이 헤더 완전 복원.
  • 테스트: pipeline_table_roundtrip 을 2행·헤더 셀 내용 단언으로 강화(기존엔 버그를 주석으로 박제하고 있었음), 표 셀 image/link 디코더 테스트 2건 행 인덱스 정정.

Changed — BREAKING (이전 wave 누락분 명시, API audit)

코드 audit 에서 발견한, 이미 발생했으나 Breaking 으로 명시되지 않았던 public API 변경 (downstream 마이그레이션 누락 방지):

  • Control::memo(content, author, date) → Control::memo(content) 로 시그니처 축소. anchor run 이 필요하면 Control::memo_with_anchor(content, anchor_runs) 사용 (Wave 12e/f).
  • Control::Memo 에서 author: String / date: String 필드 제거 → anchor_runs: Vec<Run> + metadata: MemoMetadata 로 대체. 패턴 매칭에서 해당 필드 바인딩 제거 필요 (Wave 12e/f).
  • RefType: TryFrom<u8> public trait impl 제거 (Wave 12m Phase 2). raw HWP5 바이트 → RefType 변환은 smithy-hwp5 경계 함수로 이동. downstream 의 직접 byte 변환 코드는 깨짐.
  • blueprint::style::{CharShape, PartialCharShape} 에 #[non_exhaustive] 추가 (B3). 외부 크레이트의 struct-literal 생성 차단 — CharShape 는 PartialCharShape::resolve(), PartialCharShape 는 default() + 필드 설정으로 생성. 향후 필드 추가가 더는 breaking 이 되지 않도록 한 번에 고정 (underline_shape 추가가 이미 깨뜨린 김에).
  • blueprint::style::{ParaShape, PartialParaShape, PartialStyle} 에도 #[non_exhaustive] 추가 (B 일관화) — CharShape 쌍과 동일한 style/shape 패밀리. ParaShape 는 #[derive(Default)] 추가(가산적, 합리적 기본값) + resolve() 또는 default()+필드설정 으로 생성. (성장형이나 외부 구성이 많은 core::{Section, Paragraph} 등은 builder/Default 선행이 필요해 후속 슬라이스로 보류 — backlog 기록.)

Fixed — HWP5→HWPX 채우기 “색 없음” → faceColor="none" (P1-5, 옵션 전수조사)

HWP5 는 “배경색 없음” 을 None fill 이 아니라 Color fill + background_color = 0xFFFFFFFF (Windows COLORREF null sentinel)로 인코딩한다. colorref_to_hwpx_color 가 (raw>>24)!=0 검사를 먼저 통과시켜 faceColor="#FFFFFFFF" (흰색)을 emit했고, 한컴 native 는 faceColor="none" 을 쓴다.

  • 신규 fill 전용 헬퍼 colorref_to_hwpx_fill_color: 0xFFFFFFFF → "none", 그 외는 기존 매핑. face/hatch 색에만 적용 — 테두리 선 색은 native 가 실제 #RRGGBB 를 쓰므로 공유 함수 그대로 둠 (검증 안 된 거동 도입 방지).
  • 검증: 기존 native sample-cell-diagonal (borderFill id=2 background=0xFFFFFFFF) 변환 결과가 faceColor="none" 으로 일치, 출력에서 #FFFFFFFF 완전 제거.
  • 테스트: fill_color_maps_no_color_sentinel_to_none.

Changed — 미상 무늬/그러데이션 type warning-first (P1-3/4, 옵션 전수조사)

border_fill 의 채우기 무늬·그러데이션 type 이 알 수 없는 raw 값일 때 조용히 기본값(무늬→솔리드, 그러데이션→LINEAR)으로 떨어지던 것을 ProjectionFallback 경고로 노출. (알려진 6개 무늬·4개 그러데이션 type 은 정상 carry 되며, 무늬 None=무늬 없음 은 경고 대상 아님.)

  • 신규 경고 subject: style.border_fill.fill_pattern, style.border_fill.gradation_type (각각 border_fill_id + raw 값 포함).
  • 정상 저작으론 도달하지 않는 방어적 경로라 native fixture 불필요.
  • 테스트: unknown_hatch_pattern_warns_but_known_and_none_stay_silent, unknown_gradation_type_warns_but_known_stays_silent.

Fixed — HWP5→HWPX 이미지 채우기 모드 12종 (P1-2, 옵션 전수조사)

배경 그림 채우기(<hc:imgBrush mode>)의 16개 모드 중 4개(TILE/TOTAL/CENTER/ ZOOM)만 매핑되고 나머지 12개가 무음으로 투명 처리(fill 드롭) 되던 문제. 바둑판 가로/세로, 가운데 위/아래, 좌/우 정렬 채우기가 전부 사라졌다.

  • hwp5_image_fill_mode_to_hwpx 를 16개 모드 전부로 완성. HWP5 raw 0-15 는 OWPML ImageBrushMode enum 과 동일 순서 1:1 (raw 1=TILE_HORZ_TOP, 7= CENTER_TOP, 10=LEFT_TOP, 14=RIGHT_BOTTOM, *Middle→*_CENTER 등).
  • 검증: 사용자 작성 native sample-cell-image-fill (4 family 대표 모드 TILE_HORZ_TOP/CENTER_TOP/LEFT_TOP/RIGHT_BOTTOM) 변환 결과가 한컴 native 와 일치, 변환 경고 4→0.
  • 테스트: image_fill_mode_maps_all_16_ks_x_6101_modes (전 모드 단위).

Fixed — HWP5→HWPX 쪽 번호 형식 carry (P0-3, 옵션 전수조사)

pgnp(쪽 번호 위치) 컨트롤의 번호 모양 바이트를 전혀 읽지 않아 모든 쪽 번호가 formatType="DIGIT" 로만 나가던 문제. 로마자/한글/알파벳 쪽 번호가 조용히 아라비아 숫자로 변환됨.

  • parse_page_number_control 이 property bits 0-7 (header_data[4]) = 번호 모양(HWPNumberShape)을 읽어 NumberFormatType::try_from 으로 매핑 (위치는 기존대로 bits 8-11 = header_data[5]). HWP5 shape 코드는 NumberFormatType 과 1:1 (0=Digit, 2=RomanCapital, …).
  • 검증: 사용자 작성 native sample-pagenu-roman (로마자 대문자) 변환 결과가 한컴 native 와 byte-identical — <hp:pageNum pos="INSIDE_TOP" formatType="ROMAN_CAPITAL" sideChar="-"/>.
  • 테스트: parse_page_number_control_reads_number_shape_from_property (단위), ..._user_sample_page_number_format_matches_native (native e2e).

Fixed — HWP5→HWPX 문단 번호매기기 형식 5종 (P0-1, 옵션 전수조사)

문단 번호 형식(<hh:paraHead numFormat>) 디코딩 매핑이 KS X 6101 OWPML NumberType1 enum 과 어긋나 일부 형식이 손실/오변환됐다. 옵션 단위 전수조사 중 발견.

  • code 11: "HANJA_DIGIT" (OWPML 에 없는 무효 문자열) → CIRCLED_HANGUL_JAMO (원 ㄱ,ㄴ,ㄷ) 로 정정. 한자 숫자는 실제로 code 13 = IDEOGRAPH.
  • code 6 (CIRCLED_LATIN_CAPTION, 원 알파벳 대문자 Ⓐ), code 12 (HANGUL_PHONETIC 일,이,삼), code 13 (IDEOGRAPH 한자), code 14 (CIRCLED_IDEOGRAPH) 추가 — 이전엔 전부 무음으로 DIGIT 로 붕괴.
  • code 0~10 은 사용자 작성 native fixture sample-numbering-hangul(10수준, 9형식)로 byte 단위 일치 확인 — 가/나/다(HANGUL_SYLLABLE), ㄱ/ㄴ/ㄷ (HANGUL_JAMO) 등 정상 carry 회귀 잠금. (audit 의 “가/나/다가 아라비아로 나간다” 주장은 거짓이었음 — 코드 정독 한계, fixture 로 반증.)
  • 테스트: numbering_num_format_covers_full_ks_x_6101_enum (전 코드 단위), ..._user_sample_numbering_formats_match_native (native e2e 게이트).

Verified — HWP5→HWPX 빗금무늬(hatch fill) carry + gotcha #21 방향 스왑

빗금/역빗금 등 무늬 채우기(<hc:winBrush hatchStyle>)의 HWP5→HWPX carry 를 native 한컴 fixture 로 실측 검증. 코드 변경 없음 — 디코드/projection/인코더 경로는 이미 완성돼 있었고, 한컴 HWPX writer 의 빗금↔역빗금 문자열 스왑 (gotcha #21: 시각 빗금(/) → "BACK_SLASH", 시각 역빗금() → "SLASH") 도 이미 올바르게 반영돼 있었음.

  • fixture: 사용자 작성 native sample-charbg-hatch-{slash,backslash}.{hwp,hwpx} (글자 배경 무늬 — hatch fill 은 page/char/table border fill 에 공유되므로 동일 코드 경로를 검증). 우리 출력 winBrush 가 한컴 native 와 byte-identical (faceColor="#E5E5E5" hatchColor="#CA56A7" hatchStyle="BACK_SLASH"/"SLASH").
  • 회귀 게이트: ..._charbg_hatch_carries_swapped_pattern_direction.
  • Windows-deferred: 쪽 테두리/배경(페이지) 무늬 fixture 는 macOS 한글 [쪽] 메뉴에 항목 자체가 없어 작성 불가 → Windows 한글 fixture 대기 (masterPage gap C / non-chart OLE 와 동일 차단). 단 hatch fill 자체는 위에서 공유 검증됨.

Added — HWP5→HWPX 다단(multi-column / <hp:colPr>) carry

한 구역을 2단/3단 신문형으로 나눈 다단을 HWP5→HWPX 변환에서 보존; 이전엔 cold ctrl 을 마커로만 필터하고 단 개수를 항상 colCount="1" 로 하드코딩해 단 정보가 손실됐다.

  • HWP5 디코더: cold(CTRL_ID_COLUMN_DEF) ctrl payload 파싱 — [4..6] u16 property (bits 2-9 = 단 개수), [6..8] u16 단 간격(HWPUNIT). secd 사이드카 캡처 패턴 미러로 SectionResult.column_def 에 저장.
  • Projection: col_count >= 2 면 Section.column_settings = ColumnSettings::equal_columns(count, gap). 단일 단은 None 유지(인코더 기본값).
  • 인코더/Core: 변경 없음 — Section.column_settings 와 build_col_pr_xml 가 이미 multi-column emit + HWPX round-trip 지원. HWP5 leg 만 비어 있던 것.
  • 검증: 사용자 작성 sample-multicolumn.hwp (2단) 변환 결과 colPr 가 native 와 byte-identical (colCount="2" sameSz="1" sameGap="2268"), 0 warnings. 디코더 단위 테스트(ctrl_header_cold_captures_column_def) 추가. nextest/clippy clean, 한컴 시각 게이트 PASS.

Fixed — Blueprint 저작 시 밑줄 선종류(underline shape) 손실

Blueprint(YAML 템플릿 / 빌더 API)로 만든 문서의 밑줄이 선종류와 무관하게 항상 SOLID 로 나가던 문제. Core breaking (additive): CharShape / PartialCharShape 에 underline_shape 필드 추가 (strikeout_shape 는 이미 있었으나 underline_shape 만 누락돼 있었음).

  • 근본 원인: style_store.rs 의 Blueprint→HwpxCharShape 브리지가 underline_shape: UnderlineShape::Solid 를 하드코딩 — Blueprint 에 읽을 필드가 없었기 때문.
  • 수정: blueprint/style.rs 에 underline_shape 추가 (PartialCharShape 은 Option, CharShape 는 UnderlineShape 기본 Solid) + merge() / resolve() thread, 하드코딩을 cs.underline_shape 로 교체.
  • 범위 메모: HWP5→HWPX 변환과 HWPX→HWPX round-trip 경로는 이미 밑줄/취소선 12종 선종류를 완전히 carry 하고 있었음 (Foundation UnderlineShape/StrikeoutShape + HWP5 decoder + HWPX encoder/decoder 모두 완비). 이번 수정은 Blueprint 저작 경로 전용.
  • 검증: nextest 통과 (+ 비-Solid 밑줄 carry 회귀 테스트 + serde round-trip 테스트). 생성기 examples/underline_shapes.rs 산출물이 SOLID/DASH/DOT/DASH_DOT/DASH_DOT_DOT/LONG_DASH/DOUBLE_SLIM/WAVE 밑줄 + SOLID/DASH/DOUBLE_SLIM/WAVE 취소선을 distinct 하게 emit, 한컴 시각 게이트 PASS (이중선 2줄, 물결 곡선 정상 렌더링).

Added — HWP5↔HWPX TextArt(글맵시 / <hp:textart>) carry

한컴 글맵시(워프된 장식 문자)를 HWP5→HWPX 변환에서 보존 (이전엔 통째로 drop). Core breaking (additive): Control::TextArt { text, shape, font_name, font_style, align, line_spacing, char_spacing, width, height, horz_offset, vert_offset, fill_color, inst_id } 신설 + validate.rs 가 TextArt 를 도형 family 로 검증/그룹 자식 허용.

  • Wire 발견: TextArt 는 gso ShapeComponent(0x4C) 의 comp_type 판별자 "$tat" 가 ShapeTextArt(0x5A) sub-record 를 감싸는 구조 (0x5A 는 dead 가 아니라 TextArt 전용으로 사용됨 — WIRE_SPEC §10 정정). 0x5A 레이아웃: pt0..3 (32B) + BSTR text/fontName/fontStyle + u32 fontType(1=TTF)/textShape(0..54)/lineSpacing/charSpacing/align(0=LEFT) + 20B shadow tail.
  • textShape enum (55종): HWP5 정수 = 글맵시 모양 그리드 위치, HWPX 문자열은 native 출력에서 추출. 사용자가 작성한 56-textart fixture 로 전체 테이블 (PARALLELOGRAM=0 … WAVE2=17 … DOUBLE_LINE_CIRCLE=54) 을 한 번에 확정 (TEXTART_SHAPE_NAMES).
  • HWP5 디코더: Hwp5ShapeTextArt 스키마 + $tat/0x5A 를 4개 gso 컨텍스트에 ellipse 미러로 thread, classify_gso_control 에서 Hwp5Control::TextArt 조기 분류.
  • Projection → Control::TextArt (textart_shape_name/textart_align_name 으로 정수→문자열, 범위 밖이면 경고+fallback).
  • HWPX 인코더: <hp:textart> 직접 emit (renderingInfo scaMatrix = curSz/orgSz 계산, lineShape NONE, fillBrush, pt0-3, textartPr, shapeComment).
  • HWPX 디코더: <hp:textart> → Control::TextArt round-trip (HxTextArt).
  • 검증: workspace nextest 통과(+신규 textart 스키마/인코더/테이블 테스트), clippy -D warnings clean. 사용자 작성 sample-gso-textart-all.hwp (56 글맵시, 55 distinct shape) 변환 결과 textShape 55종 전부 보존, 0 drop, HWP5→HWPX→JSON round-trip 56개 유지. scaMatrix/orgSz/curSz native 일치.

Added — HWP5↔HWPX 묶음 객체(group / <hp:container>) carry (Wave A, flat groups)

한컴 “개체 묶기“로 묶인 그리기 객체(group)를 HWP5→HWPX 변환에서 보존 (이전엔 통째로 drop). Core breaking (additive): 재귀 Control::Group { children: Vec<Control>, width, height, horz_offset, vert_offset, inst_id } 신설 + validate.rs 가 비-도형 자식을 ValidationError::InvalidGroupChild 로 거부.

  • Wire 발견: container 는 별도 TagId(0x56) 가 아니라 gso ShapeComponent(0x4C) 의 comp_type 판별자 "$con" (line "$col" / rect "$rec" / ellipse "$ell" 와 동일 메커니즘). 자식은 한 단계 깊은 ShapeComponent 들이고 각자 기존 shape sub-record(0x4F/0x50/…) 보유.
  • HWP5 디코더: gso 스코프를 단일-Option flat 상태머신에서 GsoGroupBuilder/GsoChildBuilder 스택으로 재구성 (기존 table_stack 패턴, architect 리뷰 P0). 자식 geometry 는 자식 ShapeComponent common header 에서 파싱 (Hwp5ShapeComponentGeometry::parse_from_shape_component: x=[4..8], y=[8..12], w=[16..20], h=[20..24] — native 바이트 대조로 도출). Wave A 는 flat 만; 중첩 $con 은 depth cap(GSO_GROUP_MAX_DEPTH) 으로 warn+degrade.
  • HWPX 인코더: <hp:container> 재귀 emit. 자식 위치는 한컴이 <hc:transMatrix> translation(e3=x, e6=y)으로 잡으므로 그것을 설정 (identity matrix 면 전부 원점에 겹침 — 시각 검증으로 확인); <hp:offset> 도 동일 값으로 mirror, top-level <hp:sz>/<hp:pos> 는 자식에서 제거, <hp:curSz> 는 0×0 (native 일치).
  • HWPX 디코더: <hp:container> → Control::Group round-trip.
  • 검증: workspace nextest 통과, clippy clean, 변환 산출물 geometry 가 native sample-gso-group.hwpx 와 byte 일치 (offset/transMatrix/orgSz/ curSz/container sz·pos), 한컴 시각 게이트 PASS (직사각형+타원 나란히, 텍스트 “사각형”/“타원” 보존, 겹침 없음). 중첩 그룹 재귀는 Wave B.

Added — HWP5↔HWPX 중첩 묶음 객체($con-in-$con) 재귀 carry (Wave B)

Wave A 가 flat 그룹만 보존한 데 이어, 그룹 안의 그룹($con 이 $con 을 다시 품는 구조)을 재귀적으로 보존. Core 변경 없음(Control::Group 가 이미 재귀형 Vec<Control>); 4개 레이어가 재귀를 타도록 확장.

  • HWP5 디코더: GsoGroupBuilder.current_child 를 Option<GsoChildBuilder> → Option<GsoActiveChild> 로 교체. GsoActiveChild::{Leaf(GsoChildBuilder), Nested(Box<GsoGroupBuilder>)} 로 중첩 $con 은 자식 GsoGroupBuilder 를 열어 재귀 (Box 로 GsoGroupBuilder → GsoActiveChild → GsoGroupBuilder 타입 cycle 차단). GSO_GROUP_MAX_DEPTH 초과 시에만 leaf 로 degrade → Unknown(경고).
  • Projection: project_group_child 에 Hwp5Control::Group(nested) => project_group_run(...) arm 추가 (중첩 그룹 자식 재귀).
  • 레이아웃 힌트: collect_group_child_layout_hints 가 자식이 중첩 그룹이면 재귀 — 그러지 않으면 inner 그룹의 텍스트 자식 <hp:p> 가 hint 없이 emit 되어 paragraph layout hint count underflow 발생.
  • HWPX 인코더: encode_group_child_xml 에 Control::Group early-return 추가 — encode_group_to_xml 로 재귀 후 transMatrix/offset/curSz/sz·pos 를 다른 자식과 동일하게 처리 (groupLevel 은 재귀 내부에서 +1 baked).
  • HWPX 디코더: HxContainer.containers: Vec<HxContainer> 필드 + decode_container 가 중첩 container 재귀, MAX_NESTING_DEPTH(32) depth guard.
  • 검증: workspace nextest 통과, clippy clean. 변환 산출물이 native sample-gso-group-nested.hwpx 와 구조·geometry byte 일치 (groupLevel 0/1/2/2/1, inner container identity matrix, ellipse e3=1164/e6=6512, line e3=525/e6=13422, 전 자식 curSz 0×0 + top-level sz/pos 없음). 한컴 시각 게이트 PASS.

Fixed — SUMMERY/PATH 필드 빈 본문 → 한컴 “낮은 보안 수준 복구” 경고 (#120/#136)

HWP5 → HWPX 변환 시 문서정보(SUMMERY: $author/$title/$createtime/ $modifiedtime/$lastsaveby)·경로(PATH: $P$F) 필드의 <hp:fieldBegin>…<hp:fieldEnd> 본문이 빈 <hp:t/> 로 나가 한컴이 필드를 미해결로 보고 “낮은 보안 수준으로 복구” 경고 + 첫 열기 빈 placeholder 를 남기던 문제.

  • 진단: byte-diff 로 빈 본문이 유일 구조 차이임을 확인하고, 우리 출력에 native 값만 주입한 격리 파일로 한컴 실측 (경고 없음) 하여 본문 값이 원인임을 확정. 한컴 native HWPX 는 본문에 resolved 값을 캐싱하며 그 값은 HWP5 원본 BodyText/Section0 ParaText 의 FieldBegin..FieldEnd span 에 이미 존재. Wave 12n Step 6.6 이 “빈 본문이 경고를 회피한다” 던 주석은 틀렸음 — 당시 거부된 건 합성 ISO 날짜 (locale 불일치) 였지 verbatim 값이 아니었다.
  • Core (breaking, additive): Control::Field/UnknownSummery/ DateCodeField/PathField 에 display_text: String 추가 (CrossRef::display_text 와 동일 형태/의미, 빈 문자열 = 없음). ClickHere 는 빈 값 유지 (visible placeholder 는 hint_text).
  • HWP5 projection: FieldBegin..FieldEnd span 을 display_text 로 누적 (Hyperlink/CrossRef 와 동일 경로) + DoS cap (MAX_FIELD_DISPLAY_TEXT_UNITS = 4096, body 가 BSTR command cap 우회).
  • HWPX encoder/decoder: 본문에 캐싱값 emit + round-trip 복원 (디코더가 run body 텍스트를 display_text 로 환원).
  • 검증: workspace nextest 2,469 통과, 변환 산출물 field body 가 native 안정값과 일치 (날짜/경로는 출처 파일 차이만), 한컴 시각 게이트 (경고 없음 + 첫 열기 값 표시) PASS.

Fixed — HWP5→HWPX 차트 편집 시나리오 (크기 + custom title)

HWP5 → HWPX 변환본의 차트를 to-json → 값 편집 → from-json 으로 재생성하는 체인을 end-to-end 검증하며 발견된 2건:

  • 차트 표시 크기: projection 이 ShapeComponentOle extent (내부 캔버스, 상수 7200×7200) 를 hp:sz 로 사용해 변환 차트가 ≈2.5cm 로 축소되던 버그. 표시 프레임은 gso CtrlHeader geometry [16..24] (HWPUNIT) — 한컴 native HWPX 쌍의 hp:sz 32250×18750 과 byte 일치 (probe_chart_geometry 로 확정). geometry 우선 + extent fallback 으로 수정; encoder 의 orgSz/extent 7200 hardcode 는 native 와 일치라 유지. e2e 게이트에 native 크기 단언 추가.
  • custom title 형태: write_title 의 최소형 (<a:bodyPr/><a:lstStyle/> + 빈 <a:rPr lang="ko-KR"/>) 은 표준 OOXML 로는 유효하지만 한컴 native 형태와 다름. 사용자 작성 sample-chart-title.hwpx fixture 에서 한컴의 실제 custom title 형태를 probe 해 byte-convention 미러링으로 교체: fully-attributed bodyPr, lstStyle 제거, pPr/defRPr+rPr 에 sz=1400+함초롬돋움 4종, trailing endParaRPr. 단위 게이트 write_title_mirrors_hancom_native_form 추가.

시각 검증: 변환 → 판매 값 [10,3.5,1.5,1.2]→[1,2,3,4] 편집 + 제목 추가 → 재생성본이 한컴에서 정상 크기 프레임 + 제목 + 분리형 파이 (explosion 25% 보존) 로 렌더링 확인.

Refactor (tasks #88/#89/#90/#91/#92/#94/#95) — Wave 12 누적 부채 정리

기능 영향 0 (모든 step 에서 workspace nextest 전체 통과로 검증), 코드 구조만 개선:

  • #94 CTRL_ID 상수 33개 선언 (4개 파일, 9개 중복 그룹) → smithy-hwp5/src/ctrl_ids.rs 단일 모듈 25개 canonical 상수. drift 이름 정리: SECTION_DEF→SECD, CROSSREF→FIELD_CROSSREF, FIELD_INLINE_PAGE/INLINE_AUTONUM→ATNO, INDEXMARK_INLINE→INDEXMARK.
  • #95 Wave 12i/12k 공유 split-leader / length-prefixed UTF-16 BSTR 파서를 parse_split_leader_utf16 / parse_length_prefixed_utf16 helper 로 통합 (Dutmal main/sub, IndexMark primary/secondary). Compose (12j) 는 length prefix 가 없는 다른 wire 라 의도적으로 제외.
  • #88 parse_name_subrecord 의 Option<Option<String>> → ClickHereNameSubrecord { Named / Unnamed / Malformed } named enum.
  • #89 Hwp5ClickHereControl::parse → Result<_, ClickHereParseError> (TruncatedHeader / CommandTooLong / TruncatedCommand / CommandSyntax) — DroppedControl warning 에 구체 사유 표기.
  • #90 handle_top_level_record ~425줄 → 71줄 dispatcher + 4 helper (try_intercept_memo_cluster, try_attach_clickhere_name, handle_ctrl_header, handle_eqedit_record).
  • #91 project_paragraph_with_images_structural ~225줄 → 135줄 + 3 helper (project_visible_text_segment, start_field_from_marker, drain_unconsumed_paragraph_queues).
  • #92 smithy-hwpx/encoder/section.rs 5,481줄 → 3,765줄 (잔여 중 ~2,760줄은 root 테스트 모듈). 기존 section/table.rs 선례 패턴 (mod X; + use super::*; + pub(super) fn) 으로 8개 family 모듈 신설: field(694) / section_pr(280) / header_footer(245) / memo(169) / chart(152) / picture(139) / equation(66) / typography(64). root 재수입으로 호출처/테스트 무변경, 본문 verbatim 이동. 같은 crate 내 모듈 분리라 runtime 성능 영향 0.

Docs (task #72) — HWP5 wire spec (HwpForge-internal)

Added crates/hwpforge-smithy-hwp5/HWP5_WIRE_SPEC.md — code-grounded HWP5 binary format documentation capturing Wave 12 series discoveries that the public KS X 6101 / Hancom spec omits:

  • ParaHeader [18..22] instance_id field (Wave 12p Step 1)
  • Family-aware CtrlHeader instance_id offsets (fn/en at 16, gso/tbl/eqed at 36, Wave 12p Step 5)
  • ParaShape property1 bits 25-27 are zero-based ordinal cap=6 (Wave 12p #121)
  • Style “개요 N” / “Outline N” outline-level override for levels 7~9 (Wave 12q #122)
  • Default linesegarray synthesis values for paragraphs lacking ParaLineSeg (Wave 12p #123)
  • editable attribute per FieldType (Wave 12p #124)
  • BSTR length-prefix allocation caps (MAX_*_UNITS family, Wave 12 + task #86)
  • CTRL_ID magic constant inventory across 24 entries (3 source files, prelude to task #94 consolidation)
  • SummaryInformation OLE2 PropertySet layout incl. Hancom custom PIDs 0x14/0x15 (Wave 12o)
  • Cross-reference Command 8-parameter Hancom-canonical wire (Wave 12m, ADR-004)

Each topic links back to the canonical source file/symbol so future contributors can grep this file before re-running probes.

Linked from crates/hwpforge-smithy-hwp5/AGENTS.md for discoverability.

Wave 12q (task #122) — outline level 7~9 recovery via Style “개요 N” linkage

HWP5 ParaShape.property1 bit 25-27 은 3 bits 만 표현 가능 (cap=6). 한컴 native 는 outline level 7~9 를 Style record 의 한국어 이름 “개요 N” (N-1 = HWPX level) 으로 표현. 변환 후 paraPr 의 heading_level 을 Style “개요 N” 매핑으로 override 하여 native parity 달성.

Fixed

  • hwpforge-smithy-hwp5: apply_outline_style_level_overrides() — ParaShape 변환 후 Style 테이블에서 “개요 N” / “Outline N” 패턴 매칭 → para_shape_id 가 가리키는 paraPr 의 heading_level 을 N-1 로 override. 한컴이 wire cap=6 으로 저장한 level 7/8/9 가 정확히 emit.
  • override 는 silent (warning 없음) — semantic loss 가 아닌 lossless level recovery 이므로 audit warning_count 에 영향 없음.

Added — hwpforge-smithy-hwpx

  • HwpxStyleStore::para_shape_mut() — Wave 12q level override 용 mutable accessor.

검증

  • sample-outline-9levels.hwp (사용자 작성 native fixture) 변환 결과:
    • paraPr id 2~8 level 0~6: native parity ✅
    • paraPr id 18/16/17 level 7/8/9: hp10 namespace switch wrap 없이도 한컴이 정상 인식 ✅
  • 한컴 시각 검증: 1~7수준 1./가./1)/가)/(1)/(가)/①, 8수준 ㉠, 9~10수준 marker 없음 (native 와 byte-identical 매칭).
  • Codex(architect) 검토 반영: paraPr id 자체에 의존 안 함, heading_type 이 Outline 인 paraPr 만 override.

ADR 참조: .docs/architecture/adr/ADR-006-wave12p-task122-outline-level-7-9-diagnostics.md

Wave 12p tasks #121/#123/#124 — outline/linesegarray/editable fidelity

Fixed (#121) — hwpforge-smithy-hwp5

  • Hwp5RawParaShape::heading_level() 이 saturating_sub(1) 로 wire 를 1-based 로 잘못 가정 → 0-based 그대로 사용. HWP5 spec bit 25-27 (3 bits) 가 zero-based ordinal 0~7. HWPX <hh:heading level> 도 zero-based. Codex(architect) 검토 확인.
  • 검증: native sample-outline-9levels.hwpx 의 paraPr id 2~8 level 0~6 가 우리 emit 와 완전 일치.

Fixed (#123) — hwpforge-smithy-hwp5

  • layout_hint_patch::write_linesegarray(): HWP5 source 의 ParaLineSeg (tag 0x45) record 가 없는 paragraph 도 default lineseg 로 <hp:linesegarray> 항상 emit. 이전엔 누락 → 한컴이 “낮은 보안 수준 복구” 경고.
  • Default 값 (native fixture 도출): vertsize=1000, textheight=1000, baseline=850, spacing=600, horzsize=42520 (A4 content width), flags=393216. vertpos=0 — 한컴이 paragraph 순서 따라 재계산.

Fixed (#124) — hwpforge-foundation + hwpforge-smithy-hwpx

  • FieldType::hwpx_editable() 신규 method — Author/Title → false, LastSavedBy/CreatedTime/ModifiedTime → true (한컴 native fixture 도출). 이전엔 모든 SUMMERY field 가 editable="1" 강제됨.
  • build_summery_run_xml_raw() 시그니처 확장 (editable: bool 파라미터).
  • 검증: sample-field-docsummary.hwp 변환 결과 8/8 SUMMERY field 가 native editable 값과 일치.

Wave 12p — cross-ref target element instance ID carry (BREAKING, partial)

Wave 12m 시각 검증의 잔존 이슈 (HWP5 변환본의 footnote/endnote/figure/ table/equation cross-ref 가 한컴에서 ? 표시) 해소 작업. Steps 1-4 완료, Step 5 (HWP5 → HWPX target ID 매칭 알고리즘) 는 후속 작업.

ADR 참조: .docs/architecture/adr/ADR-005-wave12p-target-element-instance-id-carry.md

Breaking — hwpforge-foundation

  • RefContentType::BookmarkName variant 부활 (Wave 12m fixup regression revert). Bookmark N2 매핑 native 일치 보정:
    • N2=0 → Page, N2=1 → Number (책갈피 본문/번호)
    • N2=2 → BookmarkName (책갈피 이름), N2=3 → UpDownPos
  • Display 는 Contents 와 BookmarkName 모두 OBJECT_TYPE_CONTENTS emit (한컴 wire 일치, 의미는 RefType + N2 wire code 에서 결정).

Breaking — hwpforge-core

  • Image 에 pub inst_id: Option<u64> 신규.
  • Table 에 pub inst_id: Option<u64> 신규.
  • Control::Equation variant 에 inst_id: Option<u64> 신규.
  • 모두 #[non_exhaustive] + Default::default() 호환이라 builder 사용 시 caller 코드 변경 불필요.

Added — hwpforge-smithy-hwp5

  • Hwp5ParaHeader.instance_id: u32 (Step 1a) — HWP5 ParaHeader [18..22] 의 u32 LE 추출. outline cross-ref target ID source.
  • Hwp5NestedSubtree.instance_id: u32 (Step 1b) — Header/Footer/Footnote/ Endnote CtrlHeader trailer.
  • Hwp5Table.instance_id: u32 (Step 1c-1) — Table CtrlHeader trailer.
  • Hwp5EquationControl.instance_id: u32 (Step 1c-2) — eqed CtrlHeader trailer.
  • Hwp5ImageControl.instance_id: u32 (Step 1c-3) — gso CtrlHeader trailer (via NestedSubtreeContext + InlineGsoContext).
  • Helper extract_ctrl_header_trailer_instance_id(&[u8]) -> u32 — 마지막 8 bytes 중 first 4 의 u32 LE.

Changed — projection / encoder

  • HWP5 projection: 6종 target element 의 instance_id 를 Core inst_id 로 통과 (0 은 unset 으로 None 정규화).
  • HWPX encoder: <hp:footNote instId="N">, <hp:endNote instId="N">, <hp:equation id="N">, <hp:tbl id="N">, <hp:pic id="N"> 모두 Core inst_id 가 있으면 사용, 없으면 sequential generate_instid fallback.

Deferred (Step 5, 후속 작업)

Probe 결과 (crates/hwpforge-smithy-hwp5/examples/probe_instance_id.rs): HWP5 binary 의 footnote/endnote CtrlHeader payload 는 자기 자신의 instance ID 를 carry 하지 않음 (한컴이 save 시 session counter 로 synthesize). cross-ref Command 는 target ID 를 carry 하지만 target element 의 ID 는 wire 에 없음.

결과:

  • Forged HWPX path (caller 가 Image::new(...).inst_id = Some(N) 설정): ✅ 정상 동작 (encoder 가 wire 에 emit)
  • HWP5 → HWPX 변환 path: ❌ cross-ref 가 여전히 ? 표시 (target element inst_id = None, encoder skip)

Step 5 알고리즘 (ADR-005 §“Step 5 Algorithm” 에 plan 상세):

  1. Section 별 cross-ref Command target_id 들을 pre-scan + RefType 별 grouping
  2. Doc 순서대로 target element 에 그 IDs 할당 (1:N dedup 처리)
  3. cross-ref Command 는 as-is (HWP5 wire 그대로)

Step 5 land 시 HWP5 conversion path 도 cross-ref 정상 동작.

신규 examples / probe

  • crates/hwpforge-smithy-hwp5/examples/probe_instance_id.rs — debug: HWP5 fixture 의 projection inst_id 확인

검증

  • cargo nextest run --workspace --all-features: 2,463 passed, 2 skipped
  • cargo clippy --workspace --all-features --all-targets -D warnings: clean
  • 한컴 시각 검증: forged path 만 동작 확인 (HWP5 conversion path 는 Step 5 대기)

Wave 12m fixup — fieldid %xrf magic constant + RefContentType::BookmarkName 폐기 (BREAKING)

Wave 12m Phase 2 시각 검증 (한컴오피스 한글) 에서 cross-ref body 가 ? 로 표시되는 회귀 발견 → 2단계 research (1차 + 재검증) 후 두 결정적 차이 fix.

Breaking — hwpforge-foundation

  • RefContentType::BookmarkName variant 폐기. OWPML 표 156 은 ContentType 을 4종 (PAGE / NUMBER / CONTENTS / UPDOWNPOS) 만 정의 하며, 책갈피의 경우 Contents 가 “책갈피 내용/이름” 의미를 내포함 (spec 명시: “참조 대상의 제시 내용 또는 책갈피의 경우, 책갈피 내용”). Wave 12m Phase 2 Step 3 의 BookmarkName enum + OBJECT_TYPE_BOOKMARK_NAME emit 은 spec 외 invented string 으로 한컴 인식 실패. HWP5 N2 code 2 for Bookmark → RefContentType::Contents 로 매핑.

Fixed — hwpforge-smithy-hwpx

  • build_crossref_run_xml 의 fieldid 가 sequential counter (1_828_000_000 + N) 가 아니라 Hancom %xrf ASCII magic constant (0x25787266 = 628_650_598) 로 emit. native HanCom-authored HWPX 11 sample 전수 분석 (n=11) 에서 모든 CROSSREF 가 동일 magic constant 사용 확인. per-instance identity 는 id (begin_id) 가 carry. HwpForge 다른 native byte-identical 검증된 field 와 일관: ClickHere=%clk, SummeryField=%smr, PathField=%pat.
  • HWP5 projection boundary decode_hwp5_crossref_content_type: Bookmark
    • N2=2 → Contents (이전 BookmarkName). 의미는 RefType-상대적 (“책갈피 이름”).
  • HWPX encoder reverse boundary ref_content_type_wire_code: Contents → 모든 RefType 에서 N2=2 (Bookmark=“책갈피 이름”, Figure/Table/Eq/ Outline=“캡션 내용”).

Regression gates

  • crossref_builders_emit_nonzero_matching_fieldid: fieldid 기댓값 1828000000 → 628650598 변경. assertion 메시지 “non-zero derived” → “%xrf magic constant” 로 강화.

시각 검증 결과 (한컴오피스 한글 직접 열기)

케이스이전현재
Bookmark+Page (Point)?1 ✅
Bookmark+Page+Hyperlink (Point)?1 ✅
Bookmark+Contents (Span)(미검증)본문 표시 ✅
Bookmark+Contents (Point)?? (정상 — 본문 없음)

Point bookmark 의 Contents reference 는 본문이 없어 한컴이 ? 표시 — 우리 wire 형식은 정상. caller 가 SpanStart/SpanEnd 책갈피를 만들면 한컴 이 범위 안 본문을 표시.

추가 발견 (별도 task)

  • <hp:endNote> / <hp:footNote> / <hp:figure> / <hp:table> 등 cross-ref target 이 될 수 있는 elements 에 instId attribute 가 emit 되지 않음. 한컴 cross-ref Command 의 ?#<id> target lookup 이 실패하여 endnote/footnote/caption cross-ref 가 ? 표시. Wave 12m 범위 밖 — task #134 로 분리.

새 examples (시각 검증용)

  • crossref_matrix.rs — RefType × ContentType 8-permutation HWPX
  • crossref_with_target.rs — Point bookmark + cross-ref 4종
  • crossref_span_bookmark.rs — Span bookmark + Contents reference 검증

검증

  • cargo nextest run --workspace --all-features: 2,463 passed, 2 skipped
  • cargo clippy --workspace --all-features --all-targets -D warnings: clean

Wave 12m Phase 2 — Control::CrossRef HWP5 leg structured upgrade (BREAKING)

HWP5 %xrf 상호참조 필드를 Control::Unknown(tag="hwp5.crossref") surrogate (Wave 12 이전 lossy 경로) 대신 typed Hwp5CrossRefControl schema → boundary functions → native Control::CrossRef 로 end-to-end carry. 사용자 시각 검증 가능한 12 한컴 native fixture (tests/fixtures/hwp5/crossref/) e2e 회귀 게이트 동반.

ADR 참조: .docs/architecture/adr/ADR-004-wave12m-crossref-structured-upgrade.md (내부 문서).

Breaking — hwpforge-foundation

  • RefType (enum) — #[repr(u8)] 제거 + #[non_exhaustive] 추가. 신규 variants: Footnote, Endnote, Outline, Unknown(u8). 기존 TryFrom<u8> impl 제거 (boundary 책임을 smithy-hwp5 로 이동). RefType 크기 1 byte → 2 bytes (variant + discriminant).
  • RefContentType (enum) — #[repr(u8)] 제거 + #[non_exhaustive] 추가. 신규 variants: BookmarkName, Unknown(u8). 동일 크기 변화.

Breaking — hwpforge-core

  • Control::CrossRef — target_name: String 필드를 type-safe target: RefTarget 로 교체. 신규 RefTarget 열거형 (Name(String) for Bookmark, SystemId(u64) for #<id> refs, Raw(String) for unparseable fallback). Control::cross_ref(...) 생성자 시그니처도 연동 변경.
  • Control::CrossRef 에 display_text: String 필드 추가. HWPX wire 는 <hp:fieldBegin> / <hp:fieldEnd> 사이 visible run 으로 display text 를 embedding 한다. HWP5 %xrf wire 는 display text 를 직접 carry 하지 않고 ParaText 본문에 풀어 두지만, projection 이 FieldBegin..FieldEnd span 을 읽어 이 필드에 채워 넣는다. 빈 문자열은 “display text 없음” 의미.
  • JSON 호환성: 의도된 clean break. 기존 target_name / surrogate JSON 도 마이그레이션 없이 변경.

Added — hwpforge-smithy-hwp5

  • schema::section::Hwp5CrossRefControl — 9-field wire-fidelity 구조 (ctrl_id, command_raw, target_raw, ref_type_code, content_type_code, hyperlink_code, param4_raw, header_flag_raw, trailer_begin_id, trailer_field_id). MAX_CROSSREF_COMMAND_UNITS = 1024 cap 으로 할당 전 차단. 13 unit tests (security gates 포함).
  • decoder::section::Hwp5Control::CrossRef(Hwp5CrossRefControl) variant
    • CTRL_ID_FIELD_CROSSREF = 0x2578_7266 (“%xrf”) dispatch. malformed payload 는 Hwp5Warning::DroppedControl{control: "crossref"} 로 떨어진다 (silent skip 금지).
  • projection:: Boundary functions: decode_hwp5_crossref_ref_type (0~6 + Unknown(u8)), decode_hwp5_crossref_content_type (RefType-relative: Bookmark=1→Contents/2→BookmarkName, 그 외는 1→Number/2→Contents, 3→UpDownPos), decode_hwp5_crossref_target (Bookmark→Name, #<u64>→SystemId, 실패→Raw).
  • semantic::Hwp5SemanticControlKind::CrossRef 신규 variant + semantic_adapter 매핑.

Added — hwpforge-smithy-hwpx

  • Reverse boundary helpers: ref_type_wire_code(&RefType) -> u8, ref_content_type_wire_code(&RefType, &RefContentType) -> u8, crossref_target_for_command(&RefTarget, &RefType) -> String.
  • build_crossref_run_xml 가 Hancom-canonical 8-parameter form (Fiexde=1 / Prop=0 / Command=?{target};N1;N2;N3;0; / RefPath=?{target}; / RefType / RefContentType / RefHyperLink / RefOpenType=HWPHYPERLINK_JUMP_CURRENTTAB) 직접 emit. 시그니처: 1번 인자 target_name: &str → target: &RefTarget (target 종류별 정확한 emit).

Changed — hwpforge-smithy-md

  • Markdown encoder 의 Control::CrossRef arm 이 display_text 가 비어 있지 않으면 그대로 emit (사용자가 본 visible 문자열). 비어 있을 때만 target.as_display() 로 fallback (Name → “bookmark1”, SystemId → “#5”) 하고 [...] anchor brackets 로 wrap.

Removed (clean break)

  • hwpforge-smithy-hwp5::projection::HWP5_CROSSREF_UNKNOWN_TAG constant ("hwp5.crossref")
  • parse_crossref_target_name, encode_hwp5_crossref_unknown_data (smithy-hwp5 projection)
  • parse_hwp5_crossref_unknown_data, Hwp5CrossRefUnknownPayload (smithy-hwpx encoder), HWP5_CROSSREF_UNKNOWN_TAG const
  • HWPX encoder 의 Control::Unknown { tag == HWP5_CROSSREF_UNKNOWN_TAG } branch — projection 이 더 이상 emit 하지 않음.
  • build_hwp5_crossref_run_xml 는 fieldid 회귀 게이트 목적으로 #[cfg(test)] 유지.

Decoder — HWPX (Wave 12m Phase 2 Step 4)

  • <hp:fieldBegin type="CROSSREF"> 가 RefPath 의 #<u64> 를 RefTarget::SystemId 로 정규화, Bookmark refs (ref_type == Bookmark) 는 RefTarget::Name, 그 외는 RefTarget::Raw fallback. display_text capture 는 후속 wave (FieldBegin/FieldEnd 사이 sibling run 누적).

Regression gates

  • 신규 (Step 2): 12 fixture parse 회귀 게이트 (schema::section::tests::crossref_parse_*) — 모든 wire 변형의 Command 파싱 검증.
  • 신규 (Step 6): hwp5_to_hwpx_crossref_12_fixture_matrix_emits_typed_control_and_canonical_wire — 12 한컴 native fixture 의 end-to-end 변환 검증 (RefType / Hancom-canonical 8-param form / fieldid 양수).

검증

  • cargo nextest run --workspace --all-features: 2,463 passed, 2 skipped (Wave 12n Step 6 baseline 2,445 → 2,463, Wave 12m Phase 2 +18 신규 게이트).
  • cargo clippy --workspace --all-features --all-targets -D warnings: clean.

Wave 12n Step 6 — %pat PATH 필드 lossless HWPX carry

Wave 12n Step 3 에서 placeholder 로 남겨두었던 Control::PathField 의 HWPX wire 매핑을 native 형식으로 완성. #120 (한컴 “낮은 보안수준 복구” 경고의 가장 강한 트리거) 해소.

Fixed — HWPX encoder

  • Control::PathField { command } arm 이 더 이상 LOSSY SUMMERY surrogate 를 emit 하지 않고 <hp:fieldBegin type="PATH" name="" editable="0" dirty="0" zorder="-1" fieldid="628121972" metaTag="">
    • <hp:stringParam name="Format">$P|$F|$P$F</hp:stringParam> 를 직접 emit. 한컴 native 와 byte-identical (단 id 카운터 제외). Body 는 empty — 한컴이 저장 시 $P / $F / $P$F 를 실제 on-disk 경로 / 파일명으로 재평가.

Added — HWPX decoder

  • <hp:fieldBegin type="PATH"> 인식 분기 신설. Format param 우선, 없으면 Command 로 fallback. PathFieldCommand::from_wire(&cmd) 로 typed variant (Path / FileName / PathAndFileName) 또는 Unknown(s) 로 carry.

Regression gates (+2, -1)

  • 신규: pathfield_emits_native_path_wire — encoder wire 형식 검증 (type="PATH" / fieldid="628121972" / editable="0" / Format 사용 / Property 미사용)
  • 신규: pathfield_roundtrip_preserves_command_lossless — PathAndFileName / Path / FileName 세 variant 모두 lossless round-trip 검증
  • 신규: pathfield_unknown_command_roundtrips_as_unknown — 비표준 $X Command 도 PathFieldCommand::Unknown("$X") 로 carry
  • 제거: 기존 lossy_roundtrip_pathfield_becomes_unknown_summery — 이제 lossless 이므로 의도 상충. lossy 사실을 확정하던 단언이 모두 lossless 단언으로 대체됨.
  • 제거: 기존 lossy_pathfield_emits_summery_with_raw_command — 마찬가지 이유.

검증:

  • workspace nextest: 2,444 → 2,445 passed + 2 skipped (+3 신규, -2 폐기)
  • end-to-end: sample-field-docsummary.hwp 변환 결과의 PATH 필드 wire 가 한컴 native sample-field-docsummary.hwpx 와 byte-identical (single field id 카운터 차이만 존재).

Note — #120 보안 경고 보조 트리거 잔존

PATH 필드 오분류는 #120 의 가장 강한 트리거이나 #121 (outline level off-by-1), #123 (linesegarray 누락), #124 (editable bit 강제) 등 다른 트리거가 잔존. Step 6 만으로 보안 경고가 완전 사라지는지는 사용자 시각 검증 (한컴 열기 후 확인) 에 의존.

Known Issues — Wave 12p/12q 종료 후 잔존

#120 — “낮은 보안 수준 복구” 경고 (HWP5 → HWPX SUMMERY 필드)

증상 (2026-06-09 시각 검증, converted-field-docsummary-wave12p-final.hwpx):

  1. 한컴에서 변환 결과 파일을 열 때 “현재의 낮은 보안 수준으로 문서에 손상을 줄 수 있는 내용을 복구하였습니다” 경고 dialog 표시.
  2. 처음 열 때 모든 SUMMERY 필드 ($author, $title, $modifiedtime, $lastsaveby, $createtime) 의 body 가 빈 placeholder ([문서 정보 시작][문서 정보 끝]) 로 표시.
  3. 사용자가 저장하면 한컴이 자체적으로 metadata 를 evaluate 해서 hanyul / 날짜 등 실제 값으로 채워 다시 표시.

Native fixture (sample-field-docsummary.hwpx) 와의 차이:

  • native: 경고 없음 + 처음 열 때부터 SUMMERY 값 표시.
  • ours: 경고 + 처음 열 때 빈 표시 → 저장 후 update.

Wire 비교 결과 (/tmp/docsum_diff 의 grep):

native 와 우리 emit 모두 <hp:fieldBegin> 와 <hp:fieldEnd> 사이의 <hp:t> body 에 실제 값을 채움 (예: native field #2 hanyul, ours field #2 hanyul). body 자체에는 차이 없음. 그럼에도 한컴이 우리 파일은 빈 placeholder 로 표시 — 한컴이 다른 구조적 차이를 트리거로 잡는 듯.

알려지지 않은 추가 트리거 (Wave 12p/12q 후 잔존):

  • <hp:fieldBegin> 의 id/fieldid/metaTag attribute 의 값/조합이 native 와 미세하게 다른가? (예: hardcoded fieldid="628321650" vs native 의 paragraph-별 다른 값)
  • ParaText (paragraph 본문 text) 의 byte sequence 가 다른가?
  • dirty="0" 이지만 한컴이 dirty 로 인식하는 다른 signal?
  • 단일 paragraph 안에 여러 SUMMERY field 가 연속될 때 우리 emit 의 paragraph 구조가 다른가? (native field #5 는 paragraph 안에 여러 <hp:t> 가 split 되어 있음)

완화책 — 후속 wave 에서 진단 필요:

  1. native sample-field-docsummary.hwpx 와 우리 emit 의 section0.xml byte-level diff 로 트리거 specific 식별.
  2. 한컴 “낮은 보안 수준” warning 의 정확한 트리거 조건 reverse-engineering.
  3. ZIP container 의 file 순서 / mimetype 처리 차이 가능성도 검토.

user impact: 사용자가 변환 결과를 저장 한 번 하면 모든 값이 올바르게 표시되므로 functional impact 는 크지 않음. 단 첫 인상이 “빈 필드 + 경고” 라 confusing — 후속 wave 에서 trigger specific 한 fix 필요.

상태: investigation deferred — 후속 wave 에서 byte-level diff 진단 필요.

#33 — Wave 5 gap C: 별도 <hm:masterPage> element carry

상태: investigation deferred — Windows 한컴 fixture 필요.

컨텍스트 (2026-06-09 macOS 한컴 fixture 진단):

macOS 한컴은 페이지 테두리 / 머리말 / 꼬리말 / 쪽 배경 등을 모두 <hp:secPr> 안에 <hp:pageBorderFill> + <hp:header> + <hp:footer> 로 inline emit. 별도 <hm:masterPage> element 는 만들지 못함 (masterPageCnt="0", 별도 MasterPage/ 디렉토리 / Contents/masterPageN.xml 파일 없음).

진단 결과: 현재 HwpForge 변환이 이미 macOS 한컴 시나리오 (페이지 테두리 + 머리말/꼬리말) 를 정확히 처리 — <hp:pageBorderFill> 3종 (BOTH/EVEN/ODD), <hp:header applyPageType="BOTH">, <hp:footer ...> 모두 native parity (id 번호만 +1 차이).

#33 의 진짜 scope 인 별도 <hm:masterPage> element (페이지별로 다른 master template 같은 advanced 기능) 는 macOS 한컴이 만들지 못해 fixture 확보 불가. Windows 한컴 환경에서 <hm:masterPage> 정의된 .hwp + native .hwpx fixture 확보 후 진단/구현 가능.

user impact 작음: 일반 사용자가 사용하는 페이지 테두리/머리말/꼬리말 시나리오는 이미 동작. advanced <hm:masterPage> 시나리오는 한컴 자체 가 흔한 use case 가 아닌 듯 (macOS 빌드에서 누락 가능).

Wave 12o-fixup — Codex(architect) Top-5 리뷰 4건 + 종료 정직성

Codex(architect) Wave 12o post-completion 리뷰에서 발견된 P0/P1 4건의 fast-follow fix + 종료 노트 정직성 보정. Wave 12o 종료 commit (8aa47f9) 은 revert/amend 하지 않음 (감사 추적성 보존).

Security (Top-1 P0)

  • hwpforge-smithy-hwp5::schema::summary_info::parse_summary_information 의 sec_start + 8 unchecked usize add 가 32-bit / wasm 타깃에서 u32-derived sec_start >= 0xFFFFFFF8 일 때 wrap 후 panicking 슬라이스 인덱스로 진입할 수 있었음. checked_add + bytes.get(..) 패턴으로 항상 Hwp5Error::RecordParse 경로 보장. 악성 .hwp 만으로 트리거 가능한 DoS 봉쇄.

Fixed (Top-2 P0 — data carry)

  • HWPX decoder 가 한컴 emit <opf:meta name="date">2026년 …</opf:meta> 값을 extras["date"] 로 carry 하지만 HWPX encoder 의 typed-collision guard 가 이를 drop 해서 HWPX → HwpForge → HWPX 1-cycle round-trip 에서 silent data loss 발생. typed-collision 리스트에서 date 만 예외 처리하고 9-slot canonical 위치에서 extras["date"] 값을 그대로 emit. creator / subject 등 진짜 typed slot 은 계속 drop.

Fixed (Top-4 P1 — API 일관성)

  • Hwp5Decoder::decode (canonical public entry) 가 decode_intermediate 에서 생성한 intermediate.metadata 를 받았지만 projection 결과 document 에 set_metadata 하지 않아 silently drop. 한 줄 추가로 decode_hwp5_with_images / hwp5_to_hwpx_bytes 와 일관성 확보.

Fixed (S3 namespace guard — P1 신규 발견)

  • hwpforge-smithy-hwpx::decoder::metadata::enforce_namespace 가 if let Some(p) = name.prefix() 패턴이어서 unprefixed <title> / <meta> (prefix 누락) 가 S3 (namespace confusion) 가드를 통과. match 표현으로 변환해서 prefix 누락 케이스도 reject. 하스타일 metadata 슬롯 shadow 가능성 차단.

Regression gates (+3 tests)

  • summary_info::section_start_past_eof_rejected — Top-1
  • decoder::metadata::date_extras_roundtrip_preserves_value — Top-2
  • decoder::metadata::unprefixed_element_inside_metadata_rejected — S3

Workspace nextest: 2,441 → 2,444 passed + 2 skipped.

Process integrity (Top-5)

  • Wave 12o 종료 commit 의 G6/G7 PASS 표기 뒤에 G2 (native byte-parity xmllint diff), G4 (HWP5→HWPX e2e 자동 게이트), G8b (저장 후 fallback 거동) 가 미평가 상태로 묻혀 있던 사실을 CLAUDE.md 종료 노트에 명시. 후속 wave 에서 자동 게이트로 승격 약속.

Visual verification example

  • 신규 crates/hwpforge-smithy-hwpx/examples/probe_date_carry_wave12o_fixup.rs — 한컴 저장본 HWPX 를 HwpForge 가 decode → encode → 재decode 한 결과 metadata 가 동일한지 자동 검증 + content.hpf 비교 명령 출력.

Wave 12o Phase 3 — HWP5 SummaryInformation OLE2 PropertySet decoder

Added — HWP5 decoder

  • New module crate::schema::summary_info parses the \x05HwpSummaryInformation OLE2 PropertySet stream (standard MS Office layout) into a populated [Metadata](Phase 0) struct.
  • PackageReader now reads the summary stream alongside the other OLE2 sub-streams. Missing or undecodable summary downgrades to Metadata::default() with a ParserFallback warning (never fatal).
  • DecodedHwp5Intermediate gains a metadata field; both decode_hwp5_with_images and hwp5_to_hwpx_bytes forward it into the projected Document via Document::set_metadata.

Wire mapping (PropertySet PIDs → Core Metadata)

  • PID 2 (TITLE) → title
  • PID 3 (SUBJECT) → subject
  • PID 4 (AUTHOR) → author
  • PID 5 (KEYWORDS) → keywords (semicolon-split)
  • PID 6 (COMMENTS) → description
  • PID 8 (LASTAUTHOR) → last_saved_by
  • PID 0x0C (LASTSAVEDTIME, VT_FILETIME) → modified (ISO 8601)
  • PID 0x0D (CREATEDTIME, VT_FILETIME) → created (ISO 8601)
  • PID 0x14 (Hancom custom date display string) → extras["date"]
  • PID 0x15 (Hancom custom app name) → extras["appname"]
  • Unknown PIDs → extras["pid_<hex>"] (preserves wire bytes)

Security (Wave 12o architect review §11.2 M3/M4 + §11.4 S2/S7)

  • M3 — slots into Schema stage (no new pipeline stage).
  • M4 — FILETIME → ISO 8601 hand-rolled (no chrono dependency added). Sub-second precision discarded to match Hancom wire (seconds). Year > 9999 and FILETIME = 0 yield None.
  • S2 — property table monotonic-offset enforcement rejects classic PropertySet payload-cycle DoS.
  • S7 — explicit UTF-16LE BOM strip from VT_LPWSTR payloads.
  • Allocation caps: per-property body ≤ 64 KiB, property count ≤ 256.

Wave 12o Phase 2 — HWPX decoder content.hpf metadata parser + XXE/DoS defenses

Added — HWPX decoder

  • New module crate::decoder::metadata parses Contents/content.hpf <opf:metadata> into the Core [Metadata](Phase 0) struct, symmetric with Phase 1’s encoder output.
  • HwpxDecoder::decode now reads metadata before constructing the Document, so a Hancom-emitted file’s $title/$author/etc. values resolve correctly on the next encode.
  • Missing or malformed content.hpf metadata downgrades to Metadata::default() (third-party HWPX authors may omit the block).

Security (Wave 12o architect review §11.1 B3 / §11.4 S1–S3)

  • <!DOCTYPE> events are rejected explicitly — defeats XXE and billion-laughs vectors before quick-xml’s default parser path.
  • CDATA inside <opf:metadata> is rejected.
  • Allocation caps: depth ≤ 16, <opf:meta> count ≤ 256, per-text body ≤ 64 KiB, attribute value ≤ 64 KiB.
  • S3 namespace confusion guard: any child of <opf:metadata> must use the opf prefix; <hp:title> and friends are rejected.
  • S1 illegal-XML-char sanitizer is applied to all decoded text before it enters Metadata.

Wave 12o Phase 1 — HWPX encoder content.hpf metadata emit

Changed (BREAKING — public API)

  • HwpxEncoder::encode now consults document.metadata() and forwards it to the package writer.
  • package::generate_content_hpf / PackageWriter::write_hwpx signatures gain a leading metadata: &Metadata parameter (internal pub(crate); not part of the public surface).

Added — wire format

  • content.hpf <opf:metadata> is always emitted with the full nine-slot Hancom byte-parity layout (title/language/creator/subject/description/lastsaveby/ CreatedDate/ModifiedDate/date/keyword) plus alphabetical extras. None slots self-close; populated slots use the content="text" attribute.
  • date always self-closes (Hancom recomputes on save — Wave 12o §11.3 Q2). keywords are joined with ; into a single element (§11.3 Q3).
  • Defense in depth: every text body flows through sanitize_xml_text (Wave 12o §11.4 S1) before escape_xml.

Wave 12o Phase 0 — Document Metadata Core breaking

Changed (BREAKING — public API)

  • core::metadata::Metadata gains three new fields and is now annotated #[non_exhaustive]. This is ADR-003 in the 0.7.0 break window. External callers can no longer construct Metadata with a struct literal; use the new Metadata::new() seed plus the chainable .with_title()/.with_author()/.with_subject()/.with_description()/ .with_last_saved_by()/.with_keywords()/.with_created()/ .with_modified()/.with_extra() builder methods.
  • New Metadata::description: Option<String> — corresponds to HWPX <opf:meta name="description">. Distinct from subject (Hancom stores them in separate slots).
  • New Metadata::last_saved_by: Option<String> — corresponds to HWPX <opf:meta name="lastsaveby"> and the $lastsaveby SUMMERY auto-field (Wave 12n). Distinct from author (creator).
  • New Metadata::extras: BTreeMap<String, String> — lossless carry slot for <opf:meta> keys not yet promoted to typed fields. Deterministic ordering for byte-stable encoder output.
  • Rationale: HwpForge-emitted content.hpf had hardcoded <opf:title/> with all other metadata fields empty, causing Hancom Office to overwrite SUMMERY auto-field values (e.g. $title) with fallback text on save. Wave 12n’s Control::Field (Author/LastSavedBy/CreatedTime/ ModifiedTime/Title) auto-fields need a populated metadata source to resolve to user-intended values. Phases 1–4 (HWPX encoder/decoder, HWP5 SummaryInformation) wire the carry end-to-end.

Phase 12l — Control::Field name carry (Core + HWPX + HWP5)

Changed (BREAKING — public API)

  • Control::Field gains a name: Option<String> field carrying the form-mode identifier (HWPX <hp:fieldBegin name="...">). Existing callers that construct Control::Field with a field-record literal must add name: None. Pattern matches with .. are unaffected. Required because the prior model silently dropped the form-mode identifier on every HWPX↔Core↔HWP5 roundtrip.
  • Control::field() convenience constructor unchanged in signature; the produced Control::Field now has name: None.
  • Display impl for Control::Field now prints name="…" only when present.
  • Serialize/Deserialize/JsonSchema representations of Control::Field gain the new field; consumers that round-trip JSON of Control::Field must accept the new key (it is optional / defaults to null).

Fixed — HWPX encoder

  • build_field_run_xml now emits the real name attribute on <hp:fieldBegin> for both CLICK_HERE and SUMMERY (automatic) fields instead of the hardcoded name="".
  • Clickhere:set:N: now uses a fixed-point computed N (self-referential UTF-16 code unit count) instead of the literal 43 that did not match the real command string length and was a wire-truth lie even before name carry.
  • build_field_run_xml signature gains name: &str.

Fixed — HWPX decoder

  • Both CLICK_HERE/automatic-field branches and the SUMMERY branch in decode_field_control now carry fb.name into Control::Field.name (with Some("") -> None normalisation, matching the indexmark/dutmal pattern).

Added — HWP5 leg

  • HWP5 %clk CtrlHeader parser (Hwp5ClickHereControl::parse) carries the press-field’s hint_text / help_text / field_unique_id from the wire’s UTF-16LE Command string. Parser is length-driven (no :-delimiter split) and UTF-16 code unit cursor based, so embedded colons in hint/help and surrogate-pair input do not corrupt decoding.
  • HWP5 0x57 (TagId::CtrlData) sub-record parser (Hwp5ClickHereControl::parse_name_subrecord) carries the form-mode name. The %clk → 0x57 pairing is held by BodyTextParserState.pending_clickhere (mirrors the eqed → EqEdit pattern used in Wave 12d) and flushed at orphan boundaries with a targeted ProjectionFallback warning.
  • Projection has a dedicated clickhere_controls queue + new ActiveField::ClickHere variant. start_active_field / finish_active_field emit a single Control::Field Run with all four metadata fields populated; the HWPX encoder rebuilds the visible placeholder from hint_text, so the span between the FIELD_BEGIN / FIELD_END markers is intentionally not double-emitted.
  • Audit semantic model gains Hwp5SemanticControlKind::ClickHere so press-field counts attribute correctly in the audit-batch output.

Fixed — HWPX encoder Command N

  • Clickhere:set:43: was a wire-truth lie (the literal 43 did not match any real-fixture command length, even before name carry). Replaced by clickhere_command_string, which uses the empirically-derived formula N = utf16_len(rest_of_command) - 1 (verified against the seven press-field instances across five fixtures). The formula is digits(N)-independent, so no fixed-point iteration is required.

Notes

  • See .docs/algorithms/2026-06-02_clickhere_field_carry.md for the full algorithm + Codex review resolution table.
  • See .docs/research/2026-06-02_clickhere_wire_dump.md for the wire layout derivation.

Phase 12 (HWP5 drawing-object carry)

Continues the Phase 11 line: HWP5 drawing objects the decoder previously skipped now carry through Core to HWPX instead of silently emptying their host paragraph. No Core or HWPX API changes — these shape variants already existed in the shared model; only the HWP5 leg was missing.

Added — Wave 12a (GSO ellipse / arc / curve)

  • Decode gso ShapeComponentEllipse (0x50) and ShapeComponentCurve (0x53) sub-records and project them to Control::Ellipse, Control::Arc, and Control::Curve. Previously these fell through to Hwp5Control::Unknown and were dropped, emptying the host paragraph.
  • 한컴 stores arcs inside the ellipse (0x50) record with arc fields set — it does not emit a separate ShapeComponentArc (0x51). An arc is now distinguished by content and carried as Control::Arc → <hp:ellipse hasArcPr="1">.
  • Classify ellipse/arc/curve in the audit semantic model (Hwp5SemanticControlKind::{Ellipse, Arc, Curve}) so source-side control counts match converted output.
  • Binary layouts confirmed empirically from 한컴 truth fixtures (sample-gso-{ellipse,arc,curve}); golden tests assert end-to-end carry.
  • Known limitation: arcs carry as the Normal arc type sized from the bounding box. Pie/chord arc types and exact arc-sweep endpoints are deferred until dedicated fixtures exist.

Added — Wave 12b (GSO connect line)

  • Carry connectors as Control::ConnectLine → <hp:connectLine> instead of demoting them to a plain <hp:line>. 한컴 stores a connector in the same ShapeComponentLine (0x4E) sub-record as a plain line; the only discriminator is the ShapeComponent (0x4C) type tag "$col" (confirmed against $rec/$ell/$cur). A conservative guard upgrades only an exact "$col" match, so plain lines are never reclassified.
  • Classify connectors in the audit semantic model (Hwp5SemanticControlKind::ConnectLine).
  • Confirmed end to end against a natively-drawn 한컴 connector fixture (sample-gso-connectline-native: two rectangles + one connector).
  • Known limitation: only a straight connector with its endpoints is carried; the source connector’s object-link references have no <hp:connectLine> representation and are dropped.

Fixed — floating ellipse/arc/curve/connect-line positioning (HWPX encoder)

  • <hp:ellipse>, <hp:curve>, and <hp:connectLine> hardcoded inline positioning (numberingType="NONE", textWrap="TOP_AND_BOTTOM", vertRelTo/horzRelTo="PARA") instead of using the shared offset-aware helpers that <hp:line>/<hp:rect> already used. A floating shape (non-zero offset) was therefore mis-anchored to the paragraph and rendered in the wrong place in 한컴. They now route through shape_position / shape_numbering_type / shape_text_wrap, so a floating shape anchors to PAPER as PICTURE/IN_FRONT_OF_TEXT (matching 한컴) while inline shapes (zero offset) are unchanged. Exposed by Wave 12b’s first floating connector.

Added — Wave 12d (equation)

  • Carry equations as Control::Equation → <hp:equation> with the HancomEQN script preserved. The eqed ctrl used to fall through to Hwp5Control::Unknown and was dropped. The decoder now recognizes the eqed ctrl, parses the script from its child HWPTAG_EQEDIT (0x58) record (UINT32 property, then a UINT16 WCHAR-count length prefix + UTF-16 script), and projects it to Control::Equation sized from the ctrl-header geometry.
  • Classify equations in the audit semantic model (Hwp5SemanticControlKind::Equation).
  • Confirmed end to end against a 한컴 equation fixture (sample-equation-basic: {a + b} over {c + d}); golden test asserts both <hp:equation> emission and verbatim <hp:script> carry.

Added — Wave 12e-Memo (memo annotation carry + body corruption fix)

  • Carry memo annotations as Control::Memo → HWPX <hp:fieldBegin type="MEMO"> with the body paragraphs in <hp:subList>. The HWP5 %unk ctrl with command "MEMO/{shapeId}/{memo_id}/{instId}" used to fall through to Hwp5Control::Unknown, so the matching HWPTAG_MEMO_LIST (0x5D) cluster’s level-2 ParaText records would fall into the body-paragraph ParaText arm and overwrite the body text — the visible “memo content replaces body text” corpus bug.
  • Decoder now recognises the %unk MEMO/... placeholder, captures the matching cluster region at the end of the section’s last body paragraph (records at level 1/2: MemoList, ListHeader, content ParaHeader, ParaText, CharShape, …), and joins clusters back to placeholders by memo_id (not by document position) during BodyTextParserState::finish. Multi-memo fixtures confirmed both happy path and id-keyed matching.
  • Projection adds an ActiveField::MemoAnchor so the anchor text inside the FieldBegin %unk MEMO / FieldEnd span flows into runs (no drop) and the memo Run is emitted at the anchor’s start position.
  • layout_hint_patch now folds memo body paragraphs into the body scope so the HWPX patcher does not underflow the paragraph-hint queue.
  • Classify memos in the audit semantic model (Hwp5SemanticControlKind::Memo).
  • Golden tests confirmed against two 한컴 fixtures: sample-memo-basic (one memo, body anchor + body content preservation) and sample-memo-multiple (two memos, id-keyed cluster matching).

Changed — Wave 12e-Memo (Core API breaking, semver-deliberate)

  • Control::Memo no longer carries author / date — the variant is now Control::Memo { content: Vec<Paragraph> }. Neither field was actually populated by any wire path: the HWPX <hp:fieldBegin type="MEMO"> only exposes MemoShapeID / MemoType parameters, and HWP5’s %unk MEMO/... command exposes no author/date metadata. Holding the fields encouraged callers to pass dummy values that never round-tripped.
  • Control::memo(content) helper drops the corresponding author/date parameters; call sites in smithy-hwpx examples (shapes_and_references, hwpx_complete_guide_parts/section2) and tests (smithy-hwpx/src/registry_bridge.rs, smithy-hwpx/src/encoder/section.rs) updated.
  • smithy-md markdown encoder emits <!-- memo: body --> instead of <!-- memo(author): body --> (the author segment was always blank in practice).

Fixed — Wave 12f (memo anchor position)

  • HWP5 stores the memo inline FieldBegin marker with extra[0..4] = %%me (0x2525_6D65), which is not the same id as the CtrlHeader ctrl_id %unk (0x2575_6E6B). Wave 12e matched only the latter, so the inline anchor was never recognised and the memo Run ended up drained at the end of the paragraph as a point-anchored field — 한컴 rendered it as 메모 대상 문장입니다.[메모 시작][필드 끝].
  • projection.rs::start_active_field now recognises both ids: the inline marker activates ActiveField::MemoAnchor, which positions the memo Run at the correct anchor offset (vis=2 in the basic fixture).

Changed — Wave 12g (Core API breaking, semver-deliberate)

  • Control::Memo gains anchor_runs: Vec<Run>. The anchor body sits inside the same <hp:run> as fieldBegin/fieldEnd in 한컴 truth fixtures; the previous Wave-12f layout placed the anchor as a separate Run before the memo and 한컴 mis-rendered the end marker as generic [필드 끝].
  • HWPX encoder now serializes memos as a flat [fieldBegin][anchor_xml][fieldEnd] <hp:run> via build_memo_anchor_xml. See .docs/algorithms/2026-06-01_memo_anchor_serialization.md for the anchor-body collapse heuristic (single <hp:t> per anchor) and why we accept the per-run char-shape fidelity loss.
  • Control::memo_with_anchor(content, anchor_runs) helper added.

Added / Changed — Wave 12h (full <hp:parameters> carry)

  • HWPX <hp:parameters> now emits the full 한컴-standard cnt="7" block (Prop, Command, ID, Number, Author, MemoShapeIDRef, CreateDateTime) plus editable="1" dirty="1" zorder="1". The pre-12h cnt="2" block (MemoShapeID + MemoType only) made 한컴 mis-classify the field — the memo body rendered correctly, but the end marker fell back to generic [필드 끝] in 조판부호 view.
  • New schema struct Hwp5MemoCommand parses the wire’s "MEMO/{shape_id}/{memo_id}/{hancom_inst_a}/{hancom_inst_b}/{author}/{terminator}" command string. Reusable wire-string utilities (parse_ctrl_header_command_string, split_slash_command) added to schema/section.rs for future %hlk / %xrf / %bmk work.
  • New Control::Memo { metadata: MemoMetadata, … } field (Core API breaking, semver-deliberate). MemoMetadata carries shape_id_ref, number, id, author, create_datetime, command — wire values flow through projection.rs verbatim.
  • CreateDateTime has no HWP5 wire source (verified via DOC_PROPERTIES 0x10, TRACKCHANGE 0x20, MemoShape 0x5E dumps). Encoder fills it with iso8601_utc_now() (std-only Howard-Hinnant civil-from-days) at write time — matching 한컴’s own behaviour on fresh saves.
  • build_memo_parameters_xml extracted as a re-usable HWPX <hp:parameters> builder for future field types that need the same structure (hyperlink / cross-reference fidelity work).

Added / Changed — Wave 12i (dutmal carry + flat-path control filter)

  • Carry dutmal (덧말) annotations as Control::Dutmal → <hp:dutmal> with main_text / sub_text / posType / option. The decoder now reads option_raw from tail[8..12] of the tdut ctrl payload (Hwp5DutmalControl); the HWPX encoder previously hard-coded option="0" and the HWPX decoder discarded the value, so any non-default option=4 한컴 fixture lost fidelity end to end. Both legs now mirror the integer verbatim. Semantics of option are intentionally not pinned — the bit/enum meaning is undocumented and produces no visible rendering difference in our truth fixture; see .docs/algorithms/2026-06-01_dutmal_carry.md.
  • New Control::Dutmal { metadata: DutmalMetadata, … } field (Core API breaking, semver-deliberate). DutmalMetadata carries option: u32 and is #[non_exhaustive] so future sz_ratio / align / style_id_ref decode work is additive.

Fixed — Wave 12i flat-path projection control_iter filter

  • project_paragraph_with_images_flat used to iterate every control in Hwp5Paragraph.controls (including the secd / cold / %bmk / %hlk / %xrf / bokm / pgnp Unknown markers that lead a first-section paragraph) when matching inline \u{FFFC} ControlRef positions to runs. Each FFFC popped the wrong control — the marker-header Unknowns returned None from project_control_run and got dropped, while the real inline controls leaked to the end-of-paragraph drain. The observable symptom on Wave 12i’s two-dutmals-with-space fixture was the body space <hp:t> </hp:t> getting pulled in front of both dutmals (한국어 韓字 → 한국어韓字). The structural projection path was unaffected because it already separated those controls into marker_headers vs. object_controls queues. The flat path now applies the same filter so its FFFC iterator only sees object controls. Any first-section paragraph that combines secd / cold with any inline shape (rect, polygon, ellipse, image, table, equation, dutmal) is covered, not just dutmal. See .docs/algorithms/2026-06-01_dutmal_carry.md (companion-fix section) for the full root-cause + rationale.

Added / Changed — Wave 12j (compose carry + char_pr_ids fidelity)

  • Carry compose (글자겹침) annotations end-to-end through the HWP5 leg — Hwp5ComposeControl schema struct, tcps CtrlHeader decode, Hwp5Control::Compose variant, and a project_compose_run that emits Control::Compose with circleType / composeType mapped from raw enum bytes to the OWPML attribute strings (14 SHAPECIRCLETYPE values + 2 COMPOSETYPE values, including the spec-typo SHAPE_REVERSAL_TIRANGLE). HWPX→Core and Core→HWPX both already existed prior to this wave; only HWP5→Core was missing and tcps was silently dropped to Hwp5Control::Unknown.
  • New Control::Compose { char_pr_ids: Vec<u32>, … } field (Core API breaking, semver-deliberate). HWPX schema fixes <hp:compose charPrCnt> at 10, but the existing encoder hard-coded all 10 slots to u32::MAX (“no override” sentinel) and the existing decoder discarded <hp:charPr prIDRef> children entirely. The new field threads the 10 LE u32 charPr IDs verbatim through Core so a HWPX → Core → HWPX round-trip preserves which slots carry a real prIDRef override (e.g. 7) vs. the u32::MAX placeholder. Encoder pads / truncates to 10 slots.
  • Compose wire layout discriminator. The tcps CtrlHeader payload has two empirically-observed forms, discriminated by the low half of properties (data[4..6] as LE u16):
    • 0x0003 (unpacked) — composeText is fully in data[8..], body trailer carries the 4 metadata bytes + 10 × u32 charPrs. 27 of 28 round-tripped variants and every native 한컴 fixture use this layout.
    • 0x0002 (packed) — composeText[0] is in properties[2..4], composeText[1..N] is at the start of the body; the body trailer layout is unchanged. Observed exclusively on the CHAR + OVERLAP combination when 한컴 saved an HWPX → HWP5. The parser detects this via properties.low == 0x0002 and prepends the packed char to the body region before decoding.
  • Any other properties.low value falls through to Hwp5Control::Unknown. No clamp; no guessing. See .docs/algorithms/2026-06-01_compose_carry.md for the full discriminator table, the shape-glyph table for properties.high, and the validation policy rationale (Codex-reviewed).
  • The packed variant was discovered only through a 14 × 2 = 28 circleType × composeType combinatorial fixture (gen_compose_variants example) — single-shape native fixtures never trigger it. The lesson is captured in .docs/learnings/2026-06-01_hwp5_ctrl_header_properties_overloaded.md: HWP5’s CtrlHeader properties word is not a generic bitfield — each ctrl_id can repurpose it for ctrl-specific data, and the same ctrl can switch layouts within a single document.

Added — Wave 12k (IndexMark carry + 0x16 inline marker promotion)

  • Carry IndexMark (찾아보기 표시) annotations end-to-end through the HWP5 leg — Hwp5IndexMarkControl schema struct, idxm CtrlHeader decode, Hwp5Control::IndexMark variant, and a project_indexmark_run that emits Control::IndexMark { primary, secondary }. The HWPX encoder + decoder + Core variant all existed prior to Wave 12k; only the HWP5 leg was missing and idxm was silently dropped to Hwp5Control::Unknown.
  • 0x16 ParaText inline-marker promotion (Codex 🟠 HIGH). Hwp5ParaText::parse used to consume the entire 0x0E..=0x16 byte range silently. Adding a CtrlHeader dispatch alone would have left every IndexMark Run drained to end-of-paragraph instead of anchored at the inline marker. The fix carves 0x16 out of the silent-consume range and promotes it to TextSegment::ControlRef only when extra[0..4] matches the LE-stored ctrl_id for idxm. The discrimination keeps every unknown 0x16 owner falling through the silent path so we don’t alias unrelated control families.
  • Targeted parser warning. Malformed idxm payloads emit Hwp5Warning::DroppedControl { control: "indexmark", … } rather than the generic UnsupportedTag(0x47) fallback — audit baselines can attribute the loss to the IndexMark code path specifically.
  • Trailer 4 bytes deliberately discarded. The trailer is 0xFFFF_FFFF on native 한컴 authoring and 0x00000000 on HWPX→ HWP5 round-trip via 한컴 save; HWPX has no corresponding attribute, so carrying it into Core would be a format leak. The 4 bytes are still required to be present — a truncated trailer means the record boundary is no longer trustworthy.
  • Empty-secondary normalization. The source Some("") for one fixture entry comes back from 한컴 as secondary_units_len = 0; the decoder reflects 한컴’s intent by returning None. HWP5 wire cannot distinguish the two states once 한컴 has saved.
  • Golden tests cover (a) the native two-IndexMark fixture, (b) the 8-IndexMark combinatorial fixture (every primary-only and primary+secondary combination + the CPU / GPU same-paragraph order regression gate that Codex specifically called out). The schema layer adds five focused unit tests (real native bytes, real multi bytes with secondary, primary_units_len == 1 edge case, primary_units_len == 0 rejection, truncated-trailer rejection). See .docs/algorithms/2026-06-02_indexmark_carry.md for the full layout table, Codex-review rationale, and validation policy notes.

Phase 11 (HWP5 → HWPX silent-gap closure)

This release closes the largest batch of “HWP5 decoder has the bytes, but projection or shared model can’t carry them” gaps measured against truth HWPX fixtures from 한컴 Office. All work is HWP5-leg only — HWPX encode/decode paths were already verified by existing golden tests.

Added — Wave 0 (audit infrastructure)

  • audit-hwp5 warning taxonomy (SilentGap, DroppedControl, …) and --strict mode so silently dropped semantics are surfaced as build signals instead of accumulating as invisible technical debt.

Added — Wave 1 (character style line family + word break)

  • Carry underline (1b) and strikeout (1c) line family (DOT, DASH, DOT_DASH, etc.) through HwpxCharShape instead of collapsing to SOLID.
  • Carry breakLatinWord = HYPHENATION (1d) through WordBreakType and HWP5 projection so HWPX breakLatinWord is no longer always KEEP_WORD.
  • Surface silent shadow decode gap as warning (1a) until carry lands.

Added — Wave 2 (paragraph layout fidelity)

  • Carry lineSpacingType = AtLeast (2a).
  • Verify all 6 alignment variants (2b), indent + pageBreakBefore (2cd), and border + shading (2e) carry against truth HWPX fixtures.

Added — Wave 3 (paragraph-level checked decode)

  • HWP5 paragraph-level paraPr.checked decode so the third location of checkable-bullet truth (the per-item checked state) is no longer lost. Closes the legacy R1 line.

Added — Wave 4 (Field / Object parity)

  • Wave 4a: carry HWP5 Rect control through Core → HWPX.
  • Wave 4b: carry HWP5 footnote control through Core → HWPX.
  • Wave 4c: carry HWP5 chart (OLE-backed BinData) end-to-end as Control::EmbeddedChart passthrough — emits Chart/chartN.xml + BinData/oleN.ole + <hp:switch> block. Closes the DroppedControl:ole_object measurement.
  • Carry HWP5 fixed-width space and non-breaking space through Core to HWPX inline text.
  • Carry HWP5 checkable bullet checkedChar + paraHead.checkable through bullet conversion.
  • Tab fidelity end-to-end: carry inline <hp:tab width / leader / type> attributes through a new RunContent::InlineText variant; HWPX encoder/decoder updated symmetrically; fill-type → leader mapping rebuilt against openhwp truth. See debug doc .docs/debug/2026-05-26_tab_fidelity_bugs.md and .docs/debug/2026-05-27_hwpx_decoder_inline_tab_attrs_lost.md.
  • Field controls: emit non-zero fieldid for CROSSREF and keep fieldBegin id within signed 32-bit range.

Added — Wave 5 (page-level features)

  • Wave 5 gap A: carry per-ctrl applyPageType (BOTH / ODD / EVEN) through HWP5 head / foot ctrl property word into multiple <hp:header> / <hp:footer> elements. Verified against sample-header-footer-odd-even truth fixture.
  • Wave 5 gap B: carry secd ctrl property bits (0/1/2/5/19) into Section.visibility.hide_first_header / footer / page_num / empty_line / master_page. Verified against sample-header-footer-hide-first truth fixture. New crates/hwpforge-smithy-hwp5/examples/probe_secd.rs reusable probe.
  • Wave 5 gap C (masterPage carry): deferred. Diagnosed as fixture-asymmetric on macOS 한컴 (truth HWPX has masterpage0.xml but the paired HWP5 has no master-page sub-records). See debug doc .docs/debug/2026-05-27_hwp5_page_features_lost.md for full probe results. Resume when a PC-한컴 fixture is available.

Fixed — Wave 6 (corpus-driven conversion robustness)

Measured by extending the audit-hwp5 signal source from synthetic fixtures to the real government-document corpus and clustering the pre-categorized conversion failures. The two real bugs recovered all 29 hwp5_convert_failed documents (plus one more from 탈락); the remaining failures are inputs that are genuinely not HWP5.

  • Schemeless hyperlink URLs (e.g. www.motie.go.kr) are normalized to http:// instead of aborting the whole conversion. The HWPX encoder previously rejected any URL outside the http:// / https:// / mailto: allowlist; explicit unsafe schemes (javascript:, data:, file:, …) are still rejected. (16 corpus docs)
  • Non-leading table header rows are demoted to normal rows (with a warning) in HWP5 projection instead of failing Core validation with NonLeadingTableHeaderRow. Real 한글 tables sometimes restate a header row mid-table; the leading header block is preserved and the stray header row is demoted so the document still converts. (14 corpus docs)
  • .hwp inputs that are actually ZIP (PK.., i.e. an HWPX saved with a .hwp extension) or a Hancom secured/DRM container (SCDS..) now fail with an actionable message instead of a raw CFB byte dump.

Changed — Breaking

  • ADR-002: hwpforge_core::Section.header: Option<HeaderFooter> → Section.headers: Vec<HeaderFooter> (and footer → footers). This changes the public Section shape so that multiple page-type-scoped headers (ODD / EVEN / BOTH) can coexist as required by HWPX wire format. Empty Vec means “no header”. JSON dump now emits a list instead of an object (or omits when empty). Patch / MD encoder / CLI / MCP consumers updated to iterate the slot. See .docs/architecture/adr/ADR-002-section-multi-header-footer-cardinality.md.

Added — Public API

  • hwpforge_core::RunContent::InlineText(InlineText) variant for mixed-content runs that include <hp:tab> (and is forward-compatible with <hp:lineBreak> / <hp:nbSpace> / <hp:fwSpace>). InlineText / InlineSegment / InlineTabAttr are #[non_exhaustive].
  • Section constructors continue to be new() / with_paragraphs(); push headers/footers via section.headers.push(HeaderFooter::…).

Migration

  • Replace section.header = Some(hf) with section.headers.push(hf).
  • Replace section.header.as_ref() with section.headers.first() (single-slot consumers) or iterate section.headers (general).
  • serde shape now uses headers / footers arrays; persisted JSON using the old keys must be migrated.
  • Match arms on RunContent may need a new InlineText(_) arm. Use RunContent::carries_text() / plain_text() for read-only consumers that don’t care about inline attribute fidelity.

[0.5.2] - 2026-05-13

Added

  • hwpforge CLI gains to-md lossy / lossless modes for HWPX → Markdown export choice.
  • HWP5 → HWPX projection preserves field controls and checkable state (precursor to Phase 11 Wave 3/4 closure).

Fixed

  • Dependency hygiene updates (sha2, GitHub Actions pinning, suppress RUSTSEC-2026-0097 non-applicable advisory).

[0.5.1] - 2026-04-13

Added

  • HWP5 → HWPX style fidelity bridge improvements (more char/para style surface preserved end-to-end).
  • HWPX char effects: preserve emboss, engrave, superscript, subscript (also covered later under Wave 1 audit).

Fixed

  • Warn on conflicting vertical-position bits instead of silently normalizing.
  • HWP5 paragraph layout hints (linesegarray, safe table height) carried to HWPX so visual diff matches truth better.
  • convert-hwp5 / audit-hwp5 / inspect summary share the same style-projection warning source.

[0.5.0] - 2026-03-20

Changed — Breaking

  • Adopt shared ordered / bullet / outline list semantics across core, blueprint, and smithy-hwpx. Markdown bridge integrated.
  • Add checkable bullet semantics (HWPX heading(type="BULLET") + bullet.checkedChar + bullet.paraHead.checkable + paraPr.checked). Markdown task lists normalize to this shared HWP semantic; ordered task lists intentionally drop numbering on the way in.
  • Tighten bullet semantics: paragraph heading_level is no longer a catch-all for list semantics (see gotcha #7 in CLAUDE.md).

Added

  • Markdown bridge: preserve task list continuation paragraphs; normalize ordered task lists.
  • Fixtures: reorganized under examples/ and tests/fixtures/.

Fixed

  • HWPX style id bridging for registry-local style ids.
  • Outline contract hardening (golden tests).

[0.4.0] - 2026-03-20

Changed

  • Promote the workspace release line to 0.4.0 for the breaking tab semantics contract in hwpforge-core and hwpforge-blueprint.
  • Add shared tab-stop semantics across the IR stack so HWPX/HWP5 codecs can preserve explicit tab definitions and paragraph tab references.

Migration

  • hwpforge_core::TabDef now includes explicit stops; downstream struct literals must initialize the new field.
  • hwpforge_blueprint::Template, ParaShape, and PartialParaShape now carry tab definition references/collections directly.
  • Consumers matching on BlueprintErrorCode should handle the new tab-related error codes explicitly.

[0.3.0] - 2026-03-19

Changed

  • Promote the workspace release line to 0.3.0 for the breaking HWPX section editing contract update.
  • Preserve-first section editing now requires preservation metadata on ExportedSection and rejects stale or legacy section exports explicitly.

[0.2.0] - 2026-03-17

Changed

  • Adopt the hwpforge-core v0.2.0 public DOM contract for richer table and image semantics.
  • Align the workspace release line and internal crate pins on 0.2.0.

Migration

  • Table, TableRow, TableCell, and Image are now #[non_exhaustive] and should be constructed via new/with_* builders instead of struct literals.
  • Table DOM now carries page-break, repeat-header, cell-spacing, border/fill, header-row, cell margin, and vertical-alignment semantics directly in hwpforge-core.
  • Image DOM now carries placement metadata directly in hwpforge-core.
  • Validation now exposes CoreErrorCode::NonLeadingTableHeaderRow; downstream code that inspects validation codes should handle it explicitly.

0.1.0 - 2026-03-06

Added

  • hwpforge: Umbrella crate with feature flags (hwpx, md, full)
  • hwpforge-foundation: Primitive types (HwpUnit, Color BGR, branded Index, enums, error codes)
  • hwpforge-core: Format-independent document model with typestate validation (Draft/Validated)
    • Document, Section, Paragraph, Run, Table, Image
    • Controls: TextBox, Footnote, Endnote, Equation, Chart (18 types)
    • Shapes: Line, Ellipse, Polygon, Arc, Curve, ConnectLine
    • References: Bookmark, CrossRef, Field, Memo, IndexMark
    • Layout: Multi-column, captions, headers/footers, page numbers, master pages
    • Annotations: Dutmal, compose characters
  • hwpforge-blueprint: YAML-based style template system
    • Template inheritance with DFS merge
    • StyleRegistry with deduplicated fonts, char shapes, para shapes
    • Built-in default template (Hancom 한컴바탕)
    • BorderFill support
  • hwpforge-smithy-hwpx: Full HWPX codec (KS X 6101)
    • Decoder: HWPX ZIP+XML -> Core Document
    • Encoder: Core Document -> HWPX ZIP+XML
    • Lossless roundtrip for all supported content
    • HancomStyleSet support (Classic/Modern/Latest)
    • 22 default styles with per-style charPr/paraPr
    • ZIP bomb defense (50MB/500MB/10k limits)
    • OOXML chart generation (18 chart types)
    • Golden fixture tests with real Hancom 한글 files
  • hwpforge-smithy-md: Markdown codec
    • GFM decoder (pulldown-cmark) with YAML frontmatter
    • Lossy encoder (readable GFM) and lossless encoder (HTML+YAML)
    • Full pipeline: MD -> Core -> HWPX verified in Hancom 한글

by Ai-Scream