Behind the work 작업의 배경

Decisions, in detail. 설계와 판단의 기록.

Product decisions, technical foundations and research: the workflow, scope choices and implemented result of each case.각 사례의 작업 흐름·범위 선택·구현 결과를 정리했습니다. 제품 판단을 먼저, 기술 기반과 연구를 뒤에 둡니다.

← Back to selected work ← 대표 작업으로
Featured 핵심

Decision case studies 의사결정 케이스 스터디

AI Writer describes the writing workflow and scope choices. Memory is a technical foundation; Harness IR is research, followed by Assessment as a supporting PoC.AI Writer의 집필 흐름과 범위 선택을 설명합니다. Memory는 기술 기반, Harness IR은 연구이며 Assessment는 보조 PoC로 뒤에 둡니다.

AI Writer System

Active development 개발 중 Writing workflow 집필 작업 흐름

Find prior settings, review the source, and bring accepted context into the next draft.이전 설정을 찾고, 원문 근거를 검토하고, 채택한 맥락으로 다음 원고를 쓰는 작업공간.

I use it for my own writing, dogfooding the product and improving the analysis display and review workflow from issues encountered during use.직접 집필에 사용하며 dogfood를 진행하고, 사용 중 발견한 불편을 바탕으로 분석 화면과 검토 흐름을 개선하고 있습니다.

User & problem대상과 문제
For long-form writers looking up prior settings and events while drafting the next section.이전 설정과 사건을 찾아 다음 집필에 활용하려는 장편 창작자를 위한 작업공간입니다.
Choice & workflow선택과 작업 흐름
Draft → inspect a proposed memory and its source → accept or reject → retrieve accepted context. Generated prose stays beside the manuscript until adopted.집필 → 기억 후보와 원문 확인 → 승인·거절 → 채택한 맥락 검색으로 연결합니다. 생성문도 채택 전에는 원고 옆의 제안으로 둡니다.
Evidence & next decision근거와 다음 판단
Drafting, generation, memory approval and retrieval run locally, with recordings of generation and source-review screens.로컬에서 집필·생성·기억 승인·검색까지 연결했습니다. 생성 제안과 기억 후보의 원문 검토 화면을 녹화해 제공합니다.
AI Writer System — a continuation being generated in the background, recorded live

Recorded workflow: the generated suggestion arrives beside the draft. The author reviews it before deciding what enters the manuscript.실제 작업 화면: 생성된 제안이 원고 옆에 도착합니다. 작성자가 검토하고 원고에 넣을 내용을 결정합니다.

7 decisions, ownership & evidence설계 판단 7개 · 직접 맡은 일 · 검증 기록
Problem 문제
The workflow centers on retrieving prior settings and reviewing their source while drafting a long manuscript. 장편 집필 중 이전 설정을 찾고 원문 근거를 확인해 다음 원고에 활용하는 작업을 중심으로 설계했습니다.
Result 결과
A writing operating system that runs end-to-end on one docker compose up , with the canonical record, the Gate, and human review in the path. 118 decision briefs · 315 verification records · 3,075 passing tests. 정본·Gate·사람 검토가 경로에 들어간 글쓰기 운영체제가 docker compose up 하나로 관통 동작합니다. 결정 브리프 118 · 검증 기록 315 · 통과 테스트 3,075.
Boundary 경계
The verification counts describe system behavior. The recorded screens show how an author reviews generated prose and memory candidates. 검증 기록은 시스템 동작의 근거입니다. 실제 화면에서는 작성자가 생성 제안과 기억 후보를 검토하는 흐름을 확인할 수 있습니다.
AI Writer System — the writing workspace

The writing workspace: drafting, generation, analysis and review around one manuscript. 원고 작업공간 — 한 편의 원고를 중심으로 생성·분석·검토가 붙어 있습니다.

AI Writer System — the candidate review screen

Candidate review: extracted memory candidates (character · event · open question) wait for a human confirm or reject before becoming memory — decision 02 as a screen. 기억 후보 검토 — 추출된 기억 후보(인물·사건·떡밥)가 기억이 되기 전 사람의 승인·거절 판정을 기다립니다. 결정 02가 화면이 된 모습입니다.

AI Writer System — memory candidates grouped by identity in the review inbox

The review inbox, grouped: nine observations of the same character arrive as one group with the judge's stated reason, so a person can approve or reject them together or one by one.묶인 검토함 — 같은 인물에 대한 관찰 9건이 판정 근거와 함께 한 그룹으로 도착하고, 사람이 그룹째로 또는 하나씩 승인·거절합니다.

AI Writer System — approving a memory candidate after checking its evidence

Review, recorded live: open the candidate, read the source quote it is anchored to, then approve — the human step that stands between an AI claim and memory. 기억 후보 검토 실제 녹화 — 후보를 열어 그것이 걸려 있는 원문 근거를 확인하고 승인합니다. AI의 주장과 기억 사이에 서 있는 사람의 단계입니다.

AI Writer System — LLM call observability dashboard on the demo project

The observability view on the demo project: calls, success rate and tokens per call site, with failed and refused calls counted in the same chart rather than filtered out.데모 프로젝트의 관측 화면 — 호출부별 호출 수·성공률·토큰을 보여 주고, 실패·거부된 호출도 걸러내지 않고 같은 차트에 셉니다.

Every LLM call site is instrumented, and failed calls are counted too. All nine call sites write a standard audit record, aggregated into a KPI view (per project, plus a global admin view). Counting only the successful calls would pin the success rate at 100% forever — an observability dashboard that can never deliver bad news is decoration, not instrumentation. LLM 호출부 전부가 계측되고, 실패한 호출도 셉니다. 9개 호출부 모두가 표준 감사 레코드를 남기고, 그것을 집계한 KPI를 화면으로 봅니다(프로젝트별 + 전역 관리자). 성공한 호출만 세면 성공률은 영구히 100%가 됩니다 — 나쁜 소식을 절대 전할 수 없는 관측 대시보드는 계측이 아니라 장식입니다.

What I owned myself 제가 직접 소유한 것

  • Contract & schema. The canonical MongoDB record, the versioned system contract, and the candidate → Gate → human review → append-only memory pipeline — designed and specified by me before any code existed. The contract is versioned ( v1.8.68 ) with every revision and its reason retained. 계약과 스키마. MongoDB 정본 레코드, 버전 관리되는 시스템 계약, candidate → Gate → 사람 검토 → append-only 기억 파이프라인 — 코드가 나오기 전에 제가 설계하고 명세했습니다. 계약은 버전 관리되며( v1.8.68 ) 모든 개정과 그 이유를 남깁니다.
  • The judgment calls that departed from the spec. Rejecting the brief's literal "OpenAI-compatible rerank adapter" because no such endpoint exists; making the reranker fall back to None while embedding falls back to a fake. Both are mine, and both are recorded as departures. 명세를 이탈한 판단. "OpenAI-호환 rerank 어댑터"라는 브리프 문언을 그런 엔드포인트가 없다는 이유로 거부한 것, 임베딩은 fake로 내리되 리랭커는 None 으로 내리게 한 비대칭 — 둘 다 제 판단이고 둘 다 이탈로 기록돼 있습니다.
  • The verification protocol. Mutation-based falsification by a separate session, two-way guards, and the rule that a missing guard blocks the verdict — I wrote the protocol, and I sign off on every verdict. 검증 프로토콜. 별도 세션의 뮤테이션 반증, 양방향 가드, 가드가 빠지면 verdict를 차단하는 규칙 — 프로토콜을 제가 썼고, 모든 판정에 제가 서명합니다.
  • What the agents wrote. Most of the implementation code and test drafts, inside those contracts. Commit authority, the decision briefs, and the accept/reject call stayed with me. 에이전트가 쓴 부분. 그 계약들 안에서의 구현 코드 대부분과 테스트 초안입니다. commit 권한, 결정 브리프, 합격·불합격 판정은 제가 가졌습니다.

01 The center is memory, not a smarter model 중심은 더 영리한 모델이 아니라 기억이다

The product hypothesis prioritizes retrieving prior settings over making generation the only centerpiece. Manuscript, settings and style can accumulate as source-linked memory; the writer selects relevant context for the next draft. Whether this beats the existing workflow needs real-use comparison. 생성만을 중심에 두기보다 이전 설정을 검색하는 경험을 우선한다는 제품 가설입니다. 원고·설정·문체를 출처가 있는 기억으로 쌓고, 다음 원고에 필요한 맥락을 가져옵니다. 기존 작업 방식보다 유용한지는 실제 사용 비교가 필요합니다.

02 AI output is not canon on arrival AI 출력은 도착 즉시 정본이 아니다

Every generated or analyzed result lands as a candidate and becomes memory only after a Gate verdict and human review. This is the load-bearing decision: the moment "whatever the AI wrote is true" is allowed, the memory itself is contaminated — and every later retrieval inherits that contamination . Memory is append-only , so a bad revision stacks a version instead of erasing the past, and every claim carries a source_ref back to its position in the original text, so the author can audit the AI's judgment down to where it came from. 생성·분석 결과는 전부 candidate 로 남고, Gate 판정과 사람의 검토를 거쳐야 기억이 됩니다. 이것이 하중을 받는 결정입니다 — "AI가 쓴 것이 곧 사실"이 되는 순간 기억 자체가 오염되고, 이후의 모든 검색이 그 오염을 물려받습니다 . 기억은 append-only 라 잘못된 갱신이 과거를 지우지 못하고 버전으로 쌓이며, 모든 주장에는 원문 위치로 되짚는 source_ref 가 붙어 작가가 AI의 판단을 출처까지 검증할 수 있습니다.

"AI output is not canon on arrival. It is a candidate until something verifies it." "AI 출력은 도착 즉시 정본이 아니다. 무언가 검증하기 전까지는 후보다."

03 The same truth-vs-cache split, deliberately repeated 같은 진실/캐시 분리를 의도적으로 반복했다

MongoDB holds the canonical record; ChromaDB (BGE-m3 vector) and Elasticsearch (nori lexical) are derived hybrid indexes maintained by an async outbox worker, and final evidence is always reloaded from the canonical record rather than trusted from the index. This is the same split I first built in the Agent Memory System below — repeated here on purpose. A structure that only works once is a coincidence; one that holds across a memory server, a product, and the system I build at work is a design I can rely on. 정본은 MongoDB가 들고, ChromaDB(BGE-m3 벡터)와 Elasticsearch(nori 어휘)는 async outbox 워커가 유지하는 파생 하이브리드 색인입니다. 최종 근거는 색인을 믿지 않고 항상 정본에서 재조회합니다. 아래 Agent Memory System 에서 처음 만든 분리를 의도적으로 반복 한 것입니다. 한 번만 통하는 구조는 우연이고, 메모리 서버·제품·회사에서 만드는 시스템까지 관통하면 그것은 믿고 쓸 수 있는 설계입니다.

04 Make the swappable parts swappable — but measure before swapping 교체할 자리는 교체 가능하게 — 다만 재기 전에는 바꾸지 않는다

Embedding and reranking became provider-neutral seams . With no endpoint configured the embedding helper falls back to a fake, but the reranker falls back to None — no reranking at all. That asymmetry is deliberate: there is no useful stand-in for "fake reranking," and shuffling at random is worse than doing nothing. The brief specified "a generic OpenAI-compatible adapter" — but OpenAI has no rerank endpoint at all . Following the wording literally would have implemented a contract that does not exist, so I built the shape the field actually shares ( POST /v1/rerank , as Cohere · Jina · Voyage · TEI expose it) and recorded the departure in the brief. A guard locks the assembly, because this failure is silent : if reranking quietly falls out, nothing looks broken — the ranking just goes back to what it was. And before any swap, an evaluation harness ( recall@k · MRR · nDCG@k ) written first, with the gold labels deliberately not filled in — and a test that fails if a labeled .jsonl ever lands in the repo. 임베딩과 리랭킹을 provider-neutral seam 으로 뺐습니다. 엔드포인트가 없으면 임베딩 헬퍼는 fake로 내려가지만, 리랭커는 None 으로 — 아예 재정렬하지 않습니다. 이 비대칭은 의도한 것입니다. 리랭킹에는 "가짜 재정렬"이라는 쓸모 있는 대체물이 없고, 무작위로 섞는 것은 no-op보다 나쁘기 때문입니다. 브리프의 문언은 "generic OpenAI-호환 어댑터"였지만 OpenAI에는 rerank 엔드포인트가 없습니다 . 글자를 그대로 따랐으면 존재하지 않는 계약 을 구현할 뻔했고, 그래서 이 분야가 실제로 공유하는 형태( POST /v1/rerank — Cohere · Jina · Voyage · TEI가 같은 모양)를 구현하고 그 이탈을 브리프에 적었습니다. 조립은 가드로 잠급니다. 이 실패는 조용하기 때문입니다 — 리랭킹이 조용히 빠져도 비는 것이 없고, 순위만 예전으로 돌아갑니다. 그리고 교체하기 전에 평가 하네스( recall@k · MRR · nDCG@k )를 먼저 썼습니다. 정답은 일부러 채우지 않았고 , 라벨링된 .jsonl 이 저장소에 들어오면 실패하는 셀을 함께 뒀습니다.

"Owning a harness is not the same as having evaluated anything." "하네스가 있다는 사실은 '평가했다'가 아니다."

05 Write the decision brief before the code — and stop 코드보다 결정 브리프를 먼저 쓰고, 멈춘다

Choices that cannot be quietly reversed later — architecture, contract literals, policy — are raised as a decision brief with an options table (option / description / upside / downside), a recommendation, and an explicit list of what I am deferring. Then implementation stops until the owner decides; guessing is not allowed. 118 decision briefs exist so far — the trail is public. The cost is real — work halts — and that is the point: a silently chosen default is a decision nobody remembers making, and those are the ones that cannot be undone later. 조용히 고르면 나중에 되돌릴 수 없는 선택 — 아키텍처·계약 리터럴·정책 — 은 선택지 표(선택지·설명·장점·단점)와 추천, 그리고 유예할 항목을 명시한 결정 브리프 로 올립니다. 그리고 오너가 결정할 때까지 구현을 멈춥니다 . 추측 구현은 금지입니다. 현재까지 결정 브리프 118개 — 기록은 공개돼 있습니다. 작업이 멈추는 비용은 실제로 발생하고, 그게 요점입니다 — 조용히 골라진 기본값은 아무도 내린 기억이 없는 결정이고, 나중에 되돌릴 수 없는 것은 바로 그런 것들입니다.

"A silently chosen default is a decision nobody remembers making." "조용히 골라진 기본값은, 아무도 내린 기억이 없는 결정이다."

06 Verification that is allowed to fail 실패가 허용되는 검증

After implementation, a different session tries to falsify the work by mutation — reverting the fix to confirm the regression guard fails again. Guards must fail in both directions: reintroducing the original bug, and over-correcting into a valid case. 315 independent verification records over 73 dated days , sitting on a suite of 3,075 passing tests (4,239 subtests) . The verdict split is 216 pass · 94 conditional · 5 outright fail — Snapshot counts; verdict proportions do not establish verification quality . One representative catch: an enforcement recorded as "complete" was true at the compose-file level but false at runtime , because already-created containers still carried the old port mapping — found by docker ps , not by the test suite. And those counts are not typed by hand: a test in the repo re-counts the records from disk and fails in both directions — add a record without raising the number, or delete one without lowering it, and it breaks. When the tally script's own verdict-parsing heuristic misread 5 records under one rule and 4 under another, I did not tune the heuristic until it looked right and then put it in the guard — I committed the script beside the records with its limits written at the top, and left the classification to a human. 구현 뒤에는 다른 세션 이 뮤테이션 으로 반증을 시도합니다 — 고친 것을 되돌려 회귀 가드가 다시 실패하는지 확인합니다. 가드는 양방향이어야 합니다: 원래 결함을 재현해도 실패하고, 과잉 교정으로 정상 경로를 깨도 실패해야 합니다. 73일치로 쌓인 독립 검증 기록 315건 , 그 아래 통과 테스트 3,075건(서브테스트 4,239) . 판정 분포는 합격 216 · 조건부 합격 94 · 불합격 5 로, 기록 시점의 집계이며 판정 비율만으로 검증 품질을 입증하지 않습니다 . 대표적으로 잡힌 것 하나: "시행 완료"로 기록된 항목이 compose 파일 수준에서는 참, 런타임에서는 거짓 이었습니다. 이미 만들어진 컨테이너가 옛 포트 매핑을 그대로 들고 있었기 때문이고, 이를 잡은 것은 테스트 스위트가 아니라 docker ps 였습니다. 그리고 이 숫자들은 손으로 적은 값이 아닙니다 — 저장소의 테스트 가 기록을 디스크에서 다시 세고 양방향으로 실패합니다. 기록을 추가하고 숫자를 안 올려도 깨지고, 파일을 지우고 숫자를 안 내려도 깨집니다. 집계 스크립트 의 판정 파싱 휴리스틱이 한 규칙에서 5건, 다른 규칙에서 4건을 오분류했을 때는 맞아 보일 때까지 휴리스틱을 다듬어 가드에 넣지 않았습니다 — 한계를 서두에 적은 채로 기록 옆에 커밋하고, 분류는 사람에게 남겼습니다.

"Three verifications in ten do not come back clean. That isn't a weak process — it's the evidence the process is real." "열 건 중 세 건이 깨끗하게 돌아오지 않는다. 절차가 허술하다는 뜻이 아니라, 절차가 진짜라는 증거다."

07 Local operation and deployment scope 로컬 운영과 배포 범위

The stack runs with docker compose up, with authentication, project ownership, quotas and administration implemented. Datastores bind to 127.0.0.1. Remote multi-host deployment remains a later scope. docker compose up으로 전체 스택이 동작하며 인증·프로젝트 소유권·quota·관리자 기능을 구현했습니다. 저장소는 127.0.0.1에 바인드하고 원격 다중 호스트 배포는 후속 범위로 남겼습니다.

How it connects: 연결점: This is the readable version of what I build at work. The company system in development ( Current work → ) uses related candidate, verification and gate boundaries for a different task — but its source is company-private. Here the whole trail is open: the briefs, the verification records, the guards, and the limits. 이것은 제가 회사에서 만드는 것의 읽을 수 있는 버전 입니다. 개발 중인 사내 시스템( 현재 실무 → )은 다른 업무에 candidate·검증·gate 경계를 적용합니다. 다만 소스가 회사 비공개입니다. 여기서는 그 자취 전부가 열려 있습니다 — 브리프, 검증 기록, 가드, 그리고 한계까지.

Agent Memory System

PoC Technical foundation 기술 기반

A shared long-term memory layer for AI assistants, exposed through MCP.MCP로 여러 AI 클라이언트가 공유하는 장기 메모리 레이어.

Decision핵심 판단
Store compacted meaning, not just logs. MongoDB is authoritative; ChromaDB is a rebuildable retrieval cache.로그 대신 압축된 의미를 저장합니다. MongoDB가 정본이고 ChromaDB는 재구축 가능한 검색 캐시입니다.
Result확인된 결과
Clients can save and recall memory, and compact it into daily, weekly, monthly, and yearly digests.클라이언트가 기억을 저장·검색하고 일·주·월·년 단위 digest로 압축하는 구조를 구현했습니다.
Limit한계
Masking is off by default. The server cannot force clients to use memory, and cache reindexing remains an operational task.가림 정책은 기본 비활성화입니다. 클라이언트의 기억 사용을 강제할 수 없으며, 캐시 재색인은 운영상 필요합니다.
5 decisions, ownership & evidence설계 판단 5개 · 직접 맡은 일 · 검증 기록
Problem 문제
An MCP PoC for saving and recalling prior context across AI-assisted work sessions. AI와 작업할 때 이전 맥락을 저장하고 다음 세션에서 검색하도록 만든 MCP PoC입니다.
Result 결과
An MCP long-term memory layer whose primary unit is compacted meaning: MongoDB authoritative, ChromaDB a rebuildable derived cache, client-driven compaction into day → week → month → year digests. 압축된 의미를 기본 단위로 삼는 MCP 장기 메모리 레이어입니다. MongoDB가 권위 저장소, ChromaDB는 재구축 가능한 파생 캐시이고, client-driven compaction이 일→주→월→년 digest를 만듭니다.
Boundary 경계
The masking policy ships disabled by default — it supports a safe policy but is not safe-by-default. And a memory server cannot force a client to actually use the memory it serves. 가림 정책은 기본값이 비활성화 입니다 — 안전 정책을 지원하지만 safe-by-default는 아닙니다. 그리고 메모리 서버는 클라이언트가 그 기억을 실제로 쓰도록 강제하지 못합니다.

What I owned myself 제가 직접 소유한 것

  • The truth/cache split. Deciding that MongoDB is authoritative and ChromaDB is a rebuildable derived cache — a storage boundary later used in AI Writer. 진실/캐시 분리. MongoDB를 권위 저장소로, ChromaDB를 재구축 가능한 파생 캐시로 둔 결정 — 이후 AI Writer에 사용한 저장 경계입니다.
  • What the server refuses to do. No hidden LLM summarization, no keyword-guessed sensitivity. Sensitivity is explicit metadata supplied by the agent; the machine coordinates, the party that understood the text decides. That refusal is a design decision, not a missing feature. 서버가 하지 않기로 한 것. 숨은 LLM 요약 없음, 키워드 기반 민감도 추측 없음. 민감도는 agent가 명시하는 메타데이터이고, 기계는 조율하며 텍스트를 이해한 주체가 결정합니다. 이 거부는 미구현이 아니라 설계 결정입니다.
  • Stating the limit. The masking policy ships disabled by default , and a memory server cannot force a client to actually use its memory. I documented both as operating boundaries rather than letting the MCP contract imply a guarantee it does not give. 한계 명시. 가림 정책은 기본값 비활성화 로 출하되고, 메모리 서버는 클라이언트에게 기억 사용을 강제할 수 없습니다. MCP 계약이 주지 않는 보장을 암시하게 두지 않고, 둘 다 운영 경계로 문서화했습니다.

01 Memory is compacted meaning, not a conversation log 기억은 대화 로그가 아니라 압축된 의미다

Most "memory" systems treat chat transcripts as the long-term unit. This one instead models a distilled, structured, time-aware memory as the primary unit: the client that understood the conversation is expected to decide what deserves saving. That is an operating contract, not a server-side ban on arbitrary text — memory quality is decided at save time. Client-driven compaction then builds day → week → month → year digests. 대부분의 '메모리' 시스템은 대화 기록을 장기 단위로 삼습니다. 이 시스템은 대신 증류된 구조적·시간 인식 memory 를 기본 단위로 모델링하고, 대화를 이해한 client가 무엇을 저장할지 결정하게 합니다. 이는 임의 텍스트를 서버가 금지한다는 뜻이 아니라 운영 계약입니다 — 기억 품질은 저장 시점에 결정됩니다. 이후 client-driven compaction이 일→주→월→년 digest를 만듭니다.

"Memory is not a log. Memory is compacted meaning." "기억은 로그가 아니다. 기억은 압축된 의미다."

02 MongoDB is authoritative; the vector store is a rebuildable cache MongoDB가 권위 저장소이고, 벡터 store는 재구성 가능한 캐시다

I split the architecture in two: MongoDB is the durable source of truth, while ChromaDB is a derived vector cache that can be rebuilt from Mongo. "Source of truth" does not mean immutable — memories can be updated or deleted. The current handlers do not fully synchronize every Chroma document inline, so reindex/rebuild remains an operational requirement; that is exactly why the cache never outranks Mongo. 아키텍처를 둘로 나눴습니다: MongoDB 는 영속적인 source of truth이고, ChromaDB 는 Mongo 기준으로 재구축하는 파생 벡터 캐시입니다. 여기서 source of truth는 불변이라는 뜻이 아닙니다 — memory는 update/delete할 수 있습니다. 현재 handler는 모든 Chroma 문서를 inline으로 완전 동기화하지 않아 reindex/rebuild가 운영상 필요하며, 바로 그래서 캐시가 Mongo보다 권위를 갖지 않습니다.

"The vector DB is a cache. The source of truth is the durable store." "벡터 DB는 캐시다. 진실은 영속 저장소에 있다."

03 The server coordinates; it never secretly summarizes 서버는 조율할 뿐, 몰래 요약하지 않는다

The server stores, selects candidates, and coordinates lifecycle — but it performs no hidden LLM summarization and never auto-guesses sensitivity from keywords. The agent sets sensitivity explicitly; meaning extraction is done by the client that actually understood the text. The machine coordinates; the understanding party decides what a memory means. 서버는 저장·후보 선택·생명주기 조율을 맡지만, 숨은 LLM 요약을 하지 않고 키워드로 민감도를 자동 추측하지 않습니다. agent가 sensitivity 를 명시하고, 의미 추출은 텍스트를 실제로 이해한 클라이언트가 합니다. 기계는 조율하고, 이해한 주체가 기억의 의미를 결정합니다.

04 Sensitivity is explicit metadata; redaction is an explicit policy 민감도는 명시적 메타데이터, 가림은 명시적 정책이다

The agent sets sensitivity explicitly at save time instead of relying on keyword guesses. Recall redacts high items only when hide_sensitive_on_recall=true ; include_sensitive=true then explicitly expands them. The honest boundary matters: the implementation ships with hiding disabled by default , so an operator must enable the policy. The mechanism is safe-capable, not safe-by-default. agent가 저장 시 sensitivity 를 키워드 추측 대신 명시합니다. Recall은 hide_sensitive_on_recall=true 일 때만 high 항목을 가리고, include_sensitive=true 로 명시적으로 펼칩니다. 중요한 실제 경계가 있습니다: 구현의 가림 정책은 기본값이 비활성화 라 운영자가 직접 켜야 합니다. 안전 정책을 지원하지만 safe-by-default는 아닙니다.

How it connects: 연결점: This truth-vs-cache split is the seed of the verified RAG I now build at work ("the vector DB is a cache; the source snapshot is the truth"), and the same division of labor runs through everything: the deterministic layer coordinates, the understanding party decides. 이 진실/캐시 분리가 지금 회사에서 만드는 검증형 RAG의 씨앗입니다("vector DB는 캐시, source snapshot이 진실"). 그리고 같은 역할 분담이 제가 만드는 모든 것을 관통합니다 — 결정론적 레이어가 조율하고, 이해한 주체가 결정합니다.

05 Named the boundary: a memory server cannot make a client remember 경계를 명시했다: 메모리 서버가 클라이언트에게 기억을 강제할 수는 없다

The MCP contract is implemented, but tool availability is not tool adoption. A client's system prompt or built-in file memory can outrank the server's prompt; a poorly rewritten recall query can also bury the right memory. I kept this as an explicit operating boundary instead of pretending the server controls agent behavior: deployment must define memory precedence in the client's own configuration , and cross-agent memories carry source_agent / source_client when provenance matters. Client guidance → MCP 계약은 구현돼 있지만, 도구가 있다는 것과 실제로 쓰인다는 것은 다릅니다. 클라이언트의 system prompt나 내장 파일 메모리가 서버 prompt보다 우선할 수 있고, 잘못 재작성된 recall query가 맞는 기억을 검색 순위 아래로 밀 수도 있습니다. 서버가 agent 행동까지 통제하는 척하지 않고 이를 명시적 운영 경계로 남겼습니다: 배포 시 클라이언트 자체 설정 에 메모리 우선순위를 정하고, 출처가 중요한 cross-agent 기억에는 source_agent / source_client 를 함께 저장합니다. 클라이언트 가이드 →

"The server can expose memory; only the client can choose to use it." "서버는 기억을 제공할 수 있지만, 그것을 쓸지는 클라이언트가 결정한다."

Harness IR

Feasibility · Mixed result 실현성 · 혼재 결과 Piece · Execution contract 조각 · 실행 계약

Test whether a provider-neutral Role IR improves structured extraction over hard-coded prompts.제공자 독립 Role IR이 하드코딩 프롬프트보다 구조화 추출에 유리한지 검증한 PoC.

Decision핵심 판단
Give the baseline the same retry budget and separate operational success from strict output quality.baseline에도 동일한 retry 예산을 주고, 실행 성공과 엄격한 출력 품질을 분리했습니다.
Result확인된 결과
One IR lowers to multiple providers; schema and evidence checks run at runtime. Hard distractors removed the clear IR advantage.하나의 IR을 여러 제공자 경로로 변환하고 schema·evidence를 검증합니다. 어려운 distractor 셋에서는 명확한 IR 우위가 사라졌습니다.
Limit한계
Mixed results, not an IR win. Different prompt shapes mean this system comparison does not isolate the causal effect of lowering.결론은 승리가 아닌 혼재입니다. prompt 형태가 달라 lowering만의 인과효과를 분리한 실험은 아닙니다.
4 decisions, ownership & evidence설계 판단 4개 · 직접 맡은 일 · 검증 기록
Problem 문제
One falsifiable question, deliberately cut out of a much larger product vision: does provider-neutral Role IR lowering beat a hardcoded prompt template for structured extraction? 훨씬 큰 제품 비전에서 의도적으로 떼어낸 반증 가능한 질문 하나입니다 — provider-neutral Role IR lowering이 하드코딩 프롬프트 템플릿보다 structured extraction에 나은가?
Result 결과
One Role IR lowers to registered OpenAI, Groq, Google GenAI/Gemma and generic OpenRouter paths, and the runtime verifies the output schema and evidence spans. Fairness was enforced against the hypothesis: the baseline got the identical retry budget. 하나의 Role IR이 등록된 OpenAI · Groq · Google GenAI/Gemma · generic OpenRouter 경로로 lowering되고, 런타임이 출력 schema와 evidence span을 검증합니다. 공정성은 가설에 불리하게 강제했습니다 — baseline에 동일한 retry 예산을 줬습니다.
Boundary 경계
Mixed, not a win. On the hard distractor set the clean IR advantage disappeared. The two prompts differ by design, so this is a controlled system-level comparison, not a string-identical ablation of lowering. 승리가 아니라 혼재입니다. 하드 distractor 셋에서 깔끔한 IR 우위가 사라졌습니다. 두 prompt는 설계상 다르므로 이는 통제된 system-level 비교이지, lowering의 인과를 재는 문자열 동일 ablation이 아닙니다.

What I owned myself 제가 직접 소유한 것

  • Experiment fairness, enforced against my own hypothesis. Only the PoC had a retry budget at first; I gave the baseline the identical budget ( match_poc ), and that alone demoted one model's "narrow win" to a tie. I also split unconditional success from strict quality so operability could not be read as quality. 내 가설에 불리하게 강제한 실험 공정성. 처음엔 PoC만 retry 예산이 있었고, baseline에 동일 예산( match_poc )을 부여했습니다. 그것만으로 한 모델의 "근소 우세"가 동급으로 재판정됐습니다. 무조건부 success 와 strict quality 도 분리해 운영성이 품질로 오인되지 않게 했습니다.
  • Lowering the conclusion when the hard set said so. The easy 8-case set favored IR; the distractor set did not. I published the uncomfortable result and rewrote the verdict from "IR wins" to mixed — and left the snapshot caveat (taken before I normalized whitespace in evidence matching) in the text instead of quietly regenerating it. 하드셋 결과에 맞춰 결론을 낮춘 것. 쉬운 8케이스 셋에선 IR이 앞섰지만 distractor 셋에선 아니었습니다. 불편한 결과를 그대로 공개하고 판정을 "IR 승리"에서 혼재로 다시 썼습니다. 스냅샷 단서(evidence 매칭 공백 정규화 이전 생성)도 조용히 재생성하지 않고 본문에 남겼습니다.
  • Refusing to name it an AI role compiler. The deterministic compiler produces a reviewable draft; it does not prove semantic equivalence. Calling it a draft generator — and keeping the human in the loop that finalizes the IR — was a naming decision I made against my own product vision. AI 역할 컴파일러라고 부르기를 거부한 것. 결정론적 컴파일러는 검토 가능한 초안을 만들 뿐 의미적 동등성을 증명하지 않습니다. 이를 초안 생성기라 부르고 IR 확정을 사람에게 남긴 것은, 제 제품 비전에 불리한 쪽으로 내린 명명 결정입니다.

01 Minimized scope to a single falsifiable hypothesis 단일 반증 가능 가설로 범위를 최소화했다

I deliberately separated the large platform vision from the POC. The runnable slice is just role_ir.yaml → lowering → one backend call → assurance , testing one claim. Proving any one of its sub-questions gives the project independent value — no need to validate everything at once. 거대한 플랫폼 비전과 POC를 의도적으로 분리했습니다. 실행 슬라이스는 role_ir.yaml → lowering → 단일 백엔드 호출 → assurance 뿐이고, 하나의 주장만 검증합니다. 하위 질문 중 하나만 증명돼도 프로젝트는 독립적 가치를 가지므로, 한 번에 전부 검증할 필요가 없습니다.

02 Enforced experiment fairness against my own hypothesis 내 가설에 불리하게, 실험 공정성을 강제했다

Initially only the POC had a retry budget. I gave the baseline the same retries ( match_poc ). That alone re-classified one model's apparent "slight edge" into a tie. I also split unconditional success from strict quality so operability couldn't be mistaken for quality. The prompts still differ by design, so this is a controlled system-level comparison—not a string-identical ablation proving a causal lowering effect. 처음엔 POC만 retry 예산을 가졌습니다. baseline에도 동일한 retry를 부여했습니다( match_poc ). 그것만으로 한 모델의 "근소 우세"가 동급으로 재판정됐습니다. 무조건부 success 와 strict quality 도 분리해, 운영성이 품질로 오인되지 않게 했습니다. 다만 두 prompt는 설계상 서로 다르므로, 이는 통제된 system-level 비교이지 lowering의 인과효과를 증명하는 문자열 동일 ablation은 아닙니다.

"Same model, output contract, assurance, case set, and retry budget—with prompt shape kept as an explicit confound." "같은 모델·출력 계약·assurance·케이스셋·retry 예산 — prompt shape 차이는 명시적 교란요인으로 남겼다."

03 Refused to cherry-pick the favorable result 유리한 결과를 cherry-pick하지 않았다

On the easier 8-case eval set IR looked slightly ahead, but on the harder sets — ones seeded with distractors designed to mislead — there was no clean IR win: false positives became a failure both paths shared. I published the uncomfortable hard-set results and reframed the conclusion: not "IR wins," but "the gain or loss splits by model, domain, and difficulty." One caveat I keep visible: the headline snapshot (2026-04-08) was generated before I normalized whitespace in evidence matching, so the top-line conclusion still holds, but some of the detailed evidence-failure counts depend on that matching rule and shouldn't be read as exact. 쉬운 8케이스 eval 셋에선 IR이 약간 앞서 보였지만, 더 어려운 셋 — 오답을 유도하도록 distractor를 심은 셋 — 에선 깔끔한 IR 승리가 없었습니다: false positive가 양쪽 경로 공통의 실패가 됐습니다. 불편한 하드셋 결과를 그대로 공개하고 결론을 재정의했습니다: "IR 승리"가 아니라 "모델·도메인·난도에 따라 이득과 손해가 갈린다." 그대로 남겨둔 단서 하나: 대표 스냅샷(2026-04-08)은 제가 evidence 매칭에서 공백을 정규화하기 전에 생성됐습니다. 그래서 top-line 결론은 여전히 유효하지만, 일부 세부 근거-실패(evidence-failure) 수치는 그 매칭 규칙에 의존하므로 정확한 값으로 읽으면 안 됩니다.

Implemented boundary: 구현된 경계: One provider-neutral Role IR lowers across registered OpenAI, Groq, Google GenAI/Gemma, and generic OpenRouter paths. The current POC validates output schema and evidence spans; policy and role-specific assurance remain part of the broader product thesis, not this runtime. 하나의 provider-neutral Role IR이 등록된 OpenAI, Groq, Google GenAI/Gemma, generic OpenRouter 경로로 lowering됩니다. 현재 PoC는 출력 schema와 evidence span을 검증하며, 정책·역할별 assurance는 이 런타임 구현이 아니라 더 큰 제품 비전에 남아 있습니다.

04 Called the compiler what it is: a draft generator 컴파일러를 있는 그대로 불렀다: 초안 생성기

The experiments above use a human-authored role_ir.yaml . The current deterministic compiler can extract an objective and schema constraints from role.md + output_schema.json , but it cannot yet prove semantic equivalence or adopt its output directly at runtime. So I did not call it an AI role compiler: it generates a reviewable draft, and a person approves the IR. The full product thesis remains larger than the slice that was actually tested. 위 실험은 사람이 작성한 role_ir.yaml 을 사용했습니다. 현재 deterministic compiler는 role.md + output_schema.json 에서 목적과 schema 제약을 추출할 수 있지만, 의미적 동등성을 증명하거나 결과를 runtime에 바로 채택하지는 못합니다. 그래서 이를 AI 역할 컴파일러라고 부르지 않았습니다: 검토 가능한 초안을 만들고, 사람이 IR을 확정합니다. 전체 제품 비전은 실제로 검증한 슬라이스보다 여전히 큽니다.

"The compiler drafts the role; it does not certify the meaning." "컴파일러는 역할 초안을 만들 뿐, 의미를 인증하지 않는다."

Assessment Spec Harness

PoC Assessment design gate 평가 설계 게이트

CI for assessment design: inspect the task specification, not the candidate.응시자 대신 과제 명세와 평가 설계를 검사하는 CI.

User & problem대상과 문제
A PoC for hiring-team designers to check mismatches between public task instructions and private scoring criteria.채용 과제 설계자를 대상으로 공개 안내와 비공개 채점 기준의 불일치를 검사하는 PoC입니다.
Scope choice범위 선택
Inspect the assessment design rather than grade applicants. Keep checks advisory and human review separate from the release gate; exclude automatic policing of grader discretion.응시자 채점 대신 평가 설계를 검사합니다. 검사 결과는 잠정 판단으로 두고 사람 검토와 배포 gate를 분리하며, 평가자의 재량을 자동 단속하는 규칙은 제외했습니다.
Evidence & next decision근거와 다음 판단
A JSON CLI and offline synthetic tests exist. Live LLM extraction and a real spec/rubric pilot remain open; those results must guide whether the scope is useful.JSON CLI와 합성 자료의 offline 검증이 있습니다. live LLM 추출과 실제 명세·채점표 파일럿은 남아 있으며, 그 결과로 현재 범위의 유용성을 판단해야 합니다.
6 decisions, ownership & evidence설계 판단 6개 · 직접 맡은 일 · 검증 기록
Problem 문제
Assessment rounds fail one level above the graders: rubric criteria that never appeared in the public spec, "optional" items that decide the outcome, recommended times that do not match the real depth of work. 채용 과제의 실패는 평가자보다 한 단계 위에서 납니다 — 공개 명세에 없던 채점 기준, 사실상 결과를 가르는 "선택" 항목, 실제 작업 깊이와 맞지 않는 권장 시간.
Result 결과
A CI that checks the assessment design, not the candidate, catching spec/rubric mismatches before applicants see them — a deterministic core plus a separate gate that reads human review, exposed as a JSON CLI contract an agent can call. 응시자가 아니라 평가 설계를 검사하는 CI입니다. 명세·채점표 불일치를 응시자가 보기 전에 잡습니다 — 결정론적 코어와, 사람의 검토를 읽는 별도 gate, 그리고 agent가 호출할 수 있는 JSON CLI 계약.
Boundary 경계
The live LLM SDK runner is still on hold. What the offline runs proved is deterministic reproducibility on valid input — not live LLM extraction quality. live LLM SDK runner는 여전히 보류입니다. offline 실행이 증명한 것은 유효 입력 위에서의 결정론적 재현성 이지, live LLM의 추출 품질이 아닙니다.

The primary caller is an AI agent; the intended human user is an assessment designer. In JSON mode ( --output json ) every command returns the same predictable envelope — what happened ( status ), an exit_code , which command ran, and suggested next_actions — and an agent can call schema --command to learn each interface on its own. A framework-neutral AgentRunner marks the seam where a real model will plug in later, but the runners that exist today are a test fixture and an offline deterministic stand-in, not live SDK integrations. 주 호출자는 AI agent이며, 목표 사용자 가설은 채용 과제 설계자입니다. JSON 모드( --output json )에서 모든 명령은 똑같이 예측 가능한 형태로 응답합니다 — 무슨 일이 있었는지( status ), 종료 코드( exit_code ), 어떤 명령이 실행됐는지( command ), 다음에 할 일 제안( next_actions ). 그리고 agent는 schema --command 를 호출해 각 인터페이스를 스스로 익힐 수 있습니다. framework-neutral AgentRunner 는 나중에 실제 모델이 꽂힐 자리를 표시하지만, 현재 존재하는 runner는 테스트 fixture와 offline deterministic stand-in이며 live SDK 연동은 아닙니다.

What I owned myself 제가 직접 소유한 것

  • The reframing. Reading one round of submissions and retrospectives and concluding the failure sat one level up — in the assessment design, not in the graders. Aiming the tool there is the decision the whole project rests on. 문제 재정의. 한 라운드의 응시·회고 자료를 정독하고, 실패가 평가자가 아니라 한 단계 위 — 평가 설계 — 에 있다고 판정한 것. 프로젝트 전체가 그 결정 위에 서 있습니다.
  • The three rules I refused. Scoring-split inspection, automatic grader-variance detection, and contradictory-rubric detection were all proposed and all rejected: they police legitimate grader discretion instead of the actual target. Drawing scope by refusal was mine. 거부한 세 규칙. 배점 분할 검사, 채점 편차 자동 감지, 모순 채점기준 검출 — 셋 다 제안됐고 셋 다 거부했습니다. 진짜 표적 대신 평가자의 정당한 재량을 단속하기 때문입니다. '거부로 범위를 긋는' 것이 제 몫이었습니다.
  • The check/gate separation and the label split. Keeping check advisory so a linter never becomes an automatic judge, and splitting a label that was quietly lying into structurally_validated vs validated after an AI audit exposed it. check/gate 분리와 라벨 분할. linter가 자동 심판자로 변질되지 않도록 check 를 잠정 판정으로 묶은 것, 그리고 AI 감사가 드러낸 뒤 조용히 거짓말하던 라벨을 structurally_validated 와 validated 로 쪼갠 것.
  • Retracting my own verdict. When an early verification found a rule's boundary check incomplete, I retracted an already-issued PASS and reissued it. Nobody made me do that. 내가 내린 판정의 철회. 초기 검증에서 한 규칙의 경계 검사가 불완전하다는 게 드러났을 때, 이미 발행한 합격 판정을 철회하고 다시 발행했습니다. 아무도 시키지 않은 일입니다.

01 Reframed the problem: judge the design, not the candidate 문제 재정의: 응시자가 아니라 설계를 평가한다

Reading one assessment round and its retrospective, I saw a systemic gap in the assessment infrastructure — recommended hours vs. real depth, "optional" items that were actually the deciding signal, grading criteria absent from the public brief. So I targeted the upstream failure, not the candidate. 한 라운드의 응시·회고 자료를 정독하니 개별 평가자의 실수가 아니라 평가 인프라 자체의 시스템적 공백 이 보였습니다 — 권장 시간 vs 실제 작업 깊이, 사실상 결정적이던 "선택" 항목, 공개 명세에 없던 채점 기준. 그래서 응시자가 아니라 그 위쪽 단계의 실패를 겨눴습니다.

"This tool does not evaluate the candidate. It evaluates the assessment design itself." "이 도구는 응시자를 평가하지 않는다. 평가 설계 자체를 평가한다."

02 Drew the scope by what I refused to build '거부'로 범위를 그었다

Three rules were proposed that I turned down on principle: checking how an evaluator distributes scoring weight, auto-detecting scoring drift, and flagging mutually contradictory criteria. Each one polices a judgment call the candidate can already see, instead of the tool's actual target — criteria that were never disclosed to the candidate at all. A PoC's default failure mode is scope creep, and every refusal keeps the tool explainable in one sentence. 제안됐지만 원칙적으로 거부한 세 규칙이 있습니다: 평가자가 배점을 어떻게 나누는지 검사하는 것, 채점 편차를 자동으로 감지하는 것, 서로 모순되는 채점 기준을 잡아내는 것. 셋 다 응시자가 이미 볼 수 있는 '평가자의 정당한 재량'을 단속할 뿐, 정작 이 도구가 노리는 표적 — 응시자에게 아예 공개되지 않은 기준 — 은 잡지 못합니다. PoC의 기본 실패 모드는 scope creep이고, 각 거부가 도구를 한 문장으로 설명 가능하게 유지합니다.

"A tool is defined as much by what it refuses to do as by what it does." "도구는 하기로 한 것만큼이나 하지 않기로 한 것으로 정의된다."

03 The machine analyzes; only the human judges 기계는 분석만, 판단은 사람만

If the deterministic checker returned pass/fail, a linter for assessment design would quietly become an automated judge . So the check step only produces provisional findings, and a separate gate step reads the human's final review before anything blocks. A blocking verdict fires for only three defects that genuinely invalidate a design: a rubric scores something the public brief never mentions, an item marked "optional" actually decides the outcome, or mandatory work earns bonus points only. 결정론적 checker가 pass/fail을 반환하면, 평가 설계의 linter 가 조용히 자동 심판자 로 변질됩니다. 그래서 check 단계는 잠정 finding만 만들고, 별도의 gate 단계가 사람의 최종 검토를 읽은 뒤에야 차단합니다. 차단 판정은 설계를 실제로 무효화하는 세 결함에만 발화합니다 — 공개 명세엔 없는 것을 채점표가 점수화하거나, "선택"이라 표시됐지만 사실상 결과를 가르는 항목이거나, 필수 작업인데 가산점으로만 인정되는 경우입니다.

04 Split a label that was quietly lying 조용히 거짓말하던 라벨을 쪼갰다

A classifier was stamping the label validated after checking only the file's shape and its audit trail — while the stricter check, that every quoted piece of evidence actually traces back to the source spec, was still unimplemented. An AI audit caught it by feeding in a run that cited a reference that didn't exist and a quote that was made up; the classifier passed it with zero errors. So I split the one label in two: structurally_validated (the shape is sound) versus validated (the evidence is genuinely grounded), so a shallow pass can never be silently promoted as the real thing. 분류기가 파일의 형식과 감사 추적(audit-trace)만 확인하고 validated 라벨을 찍고 있었습니다 — 인용된 근거가 실제로 원본 명세까지 거슬러 추적되는지 보는 더 엄격한 단계는 아직 미구현인데도요. AI 감사가 존재하지 않는 참조와 날조된 인용을 가진 run을 찔러 이를 잡아냈고, 분류기는 에러 0개로 통과시켰습니다. 그래서 라벨 하나를 둘로 분리했습니다: structurally_validated (형식은 멀쩡함)와 validated (근거까지 실제로 입증됨). 얕은 통과가 진짜인 척 조용히 승격되지 못하게 한 것입니다.

"A label that lies is worse than one more label." "거짓말하는 라벨이 라벨 하나 더 있는 것보다 나쁘다."

05 Had AI audit my own work — and allowed verdicts to be withdrawn AI에게 내 작업을 감사시키고, verdict 철회까지 허용했다

A green test suite only proves the code matches the tests, not the spec. So I made independent verification a first-class artifact, guarded in both directions — against the bug creeping back, and against valid cases getting falsely flagged — and treated a missing guard as a blocker. When an early verification record found that one rule's boundary check was incomplete, I withdrew the passing verdict and reissued it rather than quietly patching over it. green 테스트는 코드가 테스트대로 작동함만 증명할 뿐, 명세대로임을 보장하지 않습니다. 그래서 독립 검증을 1급 산출물로 삼고 양방향으로 가드를 걸었습니다 — 버그가 다시 기어들어오는 쪽과, 정상 케이스가 잘못 걸리는 쪽 모두에 대해. 누락된 가드는 차단 사유로 다뤘습니다. 초기 검증 기록에서 한 규칙의 경계 검사가 불완전하다는 걸 발견했을 때는, 조용히 덮지 않고 이미 내린 합격 verdict를 철회한 뒤 다시 발행 했습니다.

AI's role: AI의 역할: Claude Code, Codex, and Gemini were used to draft, probe, and independently audit the implementation; the CLI itself is designed for agent callers. That development workflow is separate from runtime maturity: no live model runner is implemented. I retained commit authority and final judgment. Claude Code·Codex·Gemini를 구현 초안·probe·독립 감사에 사용했고, CLI 자체도 agent caller를 위해 설계했습니다. 다만 이 개발 방식과 runtime 성숙도는 별개입니다 — live model runner는 아직 없습니다. commit 권한과 최종 판단은 제가 가졌습니다.

"I'd rather record a withdrawn verdict than ship a green bar that lies." "거짓말하는 green bar를 내보내느니 철회된 verdict를 기록하겠다."

07 Stated the limit honestly: this is still a PoC 한계를 정직하게: 아직 PoC다

The deterministic core and the review/verdict flow are implemented; live LLM SDK runners are still deferred. The offline workflow is more than a smoke test that just confirms the wiring connects: two synthetic example assignments were each run end-to-end three times, using fixed deterministic extraction plus a mock stand-in for the semantic check. Every run produced identical finding distributions and zero grounding diagnostics — no broken references, no fabricated quotes. That proves the pipeline is deterministically reproducible on valid inputs — not that a live LLM extracts well . 결정론적 코어와 review/verdict 흐름은 구현됐지만, live LLM SDK runner는 여전히 보류 상태입니다. offline workflow는 '배선이 연결되는지'만 확인하는 단순 smoke 테스트보다 깊습니다: 두 합성 예제 과제를, 고정된 결정론적 추출 + 의미 검사를 대신하는 mock으로 각각 3회 end-to-end 실행했습니다. 매번 동일한 finding 분포가 나왔고, grounding 진단(깨진 참조·날조된 인용)은 0건이었습니다. 이는 유효 입력 위에서 파이프라인이 결정론적으로 재현된다는 증명일 뿐, live LLM의 추출 품질 증명은 아닙니다 .

Working principles일하는 방식

Direction, ownership, verification.방향을 정하고, 결과를 검증합니다.

Start with the task and scope, then build with AI. Keep evidence for system behavior separate from evidence for user value.업무 문제와 범위를 먼저 정하고 AI와 구현합니다. 시스템 동작 근거와 사용자 효과의 근거를 구분합니다.

Working principles & public verification trails운영 원칙과 공개 검증 기록
How I Work 일하는 방식

Operating principles 운영 원칙

I'm not a pure-play backend or frontend engineer, and not a low-level AI researcher. My edge is the span between them — product planning and system design, plus the engineering to actually build it. I design and decide the architecture, the data contracts, and the verification boundaries; inside those, agents draft, generate candidates, and audit at speed. What does not get delegated is the call on whether a result is right — and every case study below marks exactly where that line ran. 순수 백엔드·프론트엔드 엔지니어도, 저수준 AI 연구자도 아닙니다. 제 강점은 그 사이의 폭입니다 — 제품 기획과 시스템 설계, 그리고 그것을 실제로 구현하는 엔지니어링까지. 아키텍처·데이터 계약·검증 경계는 제가 설계하고 결정하며, 그 안에서 에이전트가 초안·후보 생성·감사를 빠르게 처리합니다. 위임하지 않는 것은 "이 결과가 옳은가"라는 판정이고, 아래 모든 케이스에 그 선이 어디였는지를 표시해 뒀습니다.

I build with AI — and own the judgment. AI와 함께 만들고, 판단은 내가 책임진다.

AI agents draft code, generate candidates, and audit work at speed. I set the direction, make the engineering calls, and own whether the result is right. The capability is the pairing — and it shows in how I direct it, not in a claim. AI 에이전트가 빠르게 코드 초안·후보 생성·감사를 맡습니다. 저는 방향을 정하고 엔지니어링 판단을 내리며 결과가 옳은지를 책임집니다. 역량은 그 협업 자체이고, 주장이 아니라 그것을 어떻게 지휘하는지에서 드러납니다.

Build to validate 검증을 위해 만든다

Prototypes exist to test whether an idea can survive cost, stability, and operational friction — not to demo. 프로토타입은 데모가 아니라, 아이디어가 비용·안정성·운영 마찰을 견디는지 검증하기 위해 존재합니다.

Data-driven decisions 데이터 기반 결정

If a system can't prove its value with measurable results, I document the failure and move on. 측정 가능한 결과로 가치를 입증하지 못하면, 실패를 기록하고 다음으로 넘어갑니다.

Learn from limits 한계에서 배운다

Failed experiments, trade-offs, and dead ends are often the most useful inputs for the next design. 실패한 실험·트레이드오프·막다른 길이 다음 설계의 가장 유용한 입력입니다.

Method 방법론

How I work with AI AI를 다루는 방식

This is what AI-native engineering looks like in practice: framing the problem, directing agents, auditing their output, and keeping one source of truth. The verification records and the judgment behind each call are the evidence — not the claim. The same pattern runs through every case study above. 이것이 실제 AI-native 엔지니어링의 모습입니다 — 문제를 프레이밍하고, 에이전트를 지휘하고, 산출물을 감사하고, 단일 진실 공급원을 유지합니다. 검증 기록과 그 뒤의 판단이 주장이 아니라 증거입니다. 위의 모든 케이스 스터디에 공통으로 흐르는 패턴입니다.

Direct, don't delegate blindly 맡기지 않고 지휘한다

I frame the problem and the decision; the agent drafts, implements, and generates candidates. The judgment of "is this right?" stays with me. 문제와 결정은 제가 프레이밍하고, 에이전트는 초안·구현·후보 생성을 맡습니다. "이게 맞나?"라는 판단은 제가 가집니다.

Audit AI output independently AI 산출물을 독립 감사한다

Verification records are first-class artifacts. I have AI probe its own work with dangling references and fabricated evidence — and let it fail. 검증 기록을 1급 산출물로 둡니다. AI가 자기 작업물을 존재하지 않는 참조·날조된 근거로 찔러보게 하고, 실패하면 실패로 둡니다.

Two-way regression guards 양방향 회귀 가드

Guard against both under-strict (the bug returns) and over-strict (valid cases get flagged). A missing guard blocks the verdict. under-strict(버그 재발)와 over-strict(정상 케이스 오탐) 양쪽을 막습니다. 가드가 빠지면 verdict를 차단합니다.

One source of truth 단일 진실 공급원

With multiple agents, specs silently fork. So any contract change — even small — pays the cost of editing the canonical plan, not a side note. 여러 에이전트가 일하면 명세가 조용히 갈라집니다. 그래서 작은 계약 변경도 곁가지 노트가 아니라 정본 계획서를 편집하는 비용을 치릅니다.

It's logged, not claimed — with substance, not just a status. Across the public PoCs the trails are dated and separated by purpose: judgment, independent verification, and the daily working logs behind each. Two real excerpts: 주장이 아니라 기록입니다 — 상태값이 아니라 내용까지. 공개 PoC들에서 자취를 날짜별로, 목적별로 분리해 남깁니다 — 판단, 독립 검증, 그리고 그 뒤의 일일 워킹로그. 실제 발췌 두 개:

Verification record · Assessment · 2026-05-27 — an AI-run audit (Claude Code) of my own work, which I requested as owner: 검증 기록 · Assessment · 2026-05-27 — owner인 제가 요청해 AI(Claude Code)가 제 작업을 감사한 기록:

"This record supersedes the prior record's PASS verdict — the earlier 'default policy fallback' claim did not match the implementation." "본 기록은 직전 기록의 합격 판정을 대체한다 — 직전의 '기본 policy fallback' 주장이 실제 구현과 맞지 않았다."

Working log · Assessment · 2026-06-14 — the logs keep the owner decisions and guardrails, not just diffs: 워킹로그 · Assessment · 2026-06-14 — 로그는 diff만이 아니라 owner 결정과 가드레일까지 남깁니다:

"Produce a fully synthetic assignment — no real company/applicant data. Domain (owner decision): a synthetic PoC that mirrors the reference's mechanics without copying it." "완전 합성 과제 생성 — 실제 회사/응시자 데이터 없음. 도메인(owner 결정): 레퍼런스의 메커니즘만 모사하고 복제하지 않은 합성 PoC."

Full public trails — Assessment: decisions · verifications · working logs . Harness IR: design decisions · working logs . 전체 공개 기록 — Assessment: 판단 · 검증 · 워킹로그 . Harness IR: 설계 결정 · 워킹로그 .

Highlighted 주요

Other architecture & PoCs 그 외 아키텍처 & PoC

Pipeline

Logo Segmentation Experiment Workbench

An inspectable logo-segmentation workbench. Not a production-validated automatic extractor.실패 지점을 확인할 수 있는 로고 분할 워크벤치입니다. 프로덕션 검증을 마친 자동 추출기는 아닙니다.

Experiment record실험과 실패 분석

A /test -centered workbench that separates warp correction, anchor generation, Grounding DINO + SAM segmentation, and post-processing so each failure can be inspected. The route was earned: four SAM-prompt strategies and a vision-LLM coordinate approach failed before OCR + contour + rough vision ROI became the practical hybrid. This is not a production-validated automatic extractor ; its sharpness/noise metrics compare settings, not absolute visual quality, and cleanup cannot recover foreground that segmentation missed. Troubleshooting archive → warp 보정, anchor 생성, Grounding DINO + SAM segmentation, 후처리를 분리해 실패를 각각 관찰하는 /test 중심 워크벤치입니다. 이 경로는 설계된 게 아니라 얻어낸 것입니다: SAM prompt 전략 4개와 vision LLM 좌표 접근이 실패한 뒤 OCR + contour + 대략적 vision ROI가 현실적인 hybrid가 됐습니다. 이는 프로덕션 검증을 마친 자동 추출기가 아닙니다 . sharpness/noise 지표는 설정 비교용이지 절대적 시각 품질 판정이 아니며, 후처리는 segmentation이 놓친 전경을 복구하지 못합니다. 트러블슈팅 아카이브 →

R&D · Post-Mortems R&D · 포스트모템

Experiments & honest failures 실험과 정직한 실패

I value the decision to stop as much as the decision to ship. 출시 결정만큼이나 멈추는 결정도 중요하게 봅니다.

Project killed 프로젝트 종료

Q-PSA

Stopped in Phase 1: removing one selected layer raised PPL 3.65×, versus 1.05× with Layer Ablation.Phase 1에서 종료했습니다. 선택한 레이어 하나 제거 시 PPL이 3.65배로 상승했고, Layer Ablation은 1.05배였습니다.

Experiment record실험과 실패 분석
Pruning validation chart — Q-PSA perplexity explodes vs the Layer Ablation baseline as more layers are pruned

Q-PSA (red) vs a simple baseline — perplexity explodes as more layers are pruned. Lower is better; the gap is why I killed it. Q-PSA(빨강) vs 단순 베이스라인 — 레이어를 더 제거할수록 perplexity가 폭발. 낮을수록 좋으며, 이 격차가 프로젝트를 종료한 이유입니다.

Discrete perturbation to estimate layer importance in quantized LLMs. 양자화 LLM의 레이어 중요도를 discrete perturbation으로 추정.

Decision: Killed at Phase 1 — pruning the bottom-ranked layer raised PPL 3.65× with Q-PSA versus 1.05× with Layer Ablation, while scoring took 155 minutes vs 7 seconds (~1,300× slower). Phase 1에서 종료 — 가장 덜 중요하다고 본 레이어 하나를 제거했을 때 Q-PSA는 PPL이 3.65배 , Layer Ablation은 1.05배 였고, 점수 산출은 155분 vs 7초 (약 1,300배)였습니다.

Why it failed: perturbation sensitivity ≠ layer importance , and a gradient-free GGUF platform blocked the structural experiments needed to rescue the hypothesis. What survived was more useful than the method: Layer Ablation as a fast importance metric , a reusable in-memory GGUF perturbation pipeline, and explicit kill criteria. Full results → 실패 이유: perturbation 민감도 ≠ 레이어 중요도 , 그리고 gradient-free GGUF 플랫폼이 가설을 살리기 위한 구조 실험을 막았습니다. 방법론보다 더 쓸모 있게 남은 것은 빠른 중요도 지표인 Layer Ablation , 재사용 가능한 in-memory GGUF perturbation 파이프라인, 명시적 kill criteria였습니다. 전체 결과 →

Architectural pivot 아키텍처 피벗

Circle-WFC

Stopped treating WFC as a standalone pathfinder. Its use as a corridor proposal layer remains a research direction.WFC를 단독 경로탐색기로 보는 것을 중단했습니다. 경로 후보 제안 레이어로의 전환은 아직 연구 방향입니다.

Experiment record실험과 실패 분석
Geometry-guided WFC (ray variant) finding a 163-step path through a 100x100 maze

The ray variant finds a full 163-step path where plain WFC stalls. It demonstrates recovery on this maze, not a general connectivity guarantee. ray 변형이 plain WFC가 막히는 이 미로에서 163-step 경로를 찾아냅니다. 이 사례의 복구 결과이지, 일반적인 연결성 보장은 아닙니다.

Replacing A* pathfinding with geometry-guided Wave Function Collapse. A* 경로탐색을 geometry-guided WFC로 대체 시도.

Insight: Found the structural mismatch between local consistency and global connectivity; stopped treating it as a standalone pathfinder. local consistency와 global connectivity의 구조적 불일치를 확인하고, 단독 pathfinder로 보는 것을 중단.

The honest pivot is still a research direction, not a validated reducer : Circle-WFC could propose corridors, anchors, or road zones, while A* or JPS (jump point search) computes the final path and owns complex topology. Full result → 정직한 피벗은 아직 검증된 reducer가 아니라 후속 연구 방향 입니다. Circle-WFC가 corridor·anchor·road zone 후보를 만들고, A*나 JPS(jump point search)가 최종 경로와 복잡한 topology를 맡는 구조입니다. 전체 결과 →

Feasibility validated 실현성 검증

HW-WFC v2.9

Matched Exact DP on a synthetic benchmark, but demonstrated no practical speed advantage or production scheduling reliability.가상 벤치마크에서 Exact DP 최적값과 일치했지만, 실용적 속도 우위나 프로덕션 스케줄링 신뢰성은 입증하지 못했습니다.

Experiment record실험과 실패 분석
HW-WFC final schedule — attention block layers collapsed to optimal tile/layout states under SRAM limits, 0 backtracks

Final attention-block schedule — on the synthetic 12KB-SRAM benchmark, WFC matches exact dynamic programming. Search correctness on this case is validated; production advantage is not. 어텐션 블록의 최종 스케줄 — 가상 12KB-SRAM 벤치마크에서 WFC가 exact dynamic programming과 일치합니다. 이 사례의 탐색 정확성은 확인됐지만, 프로덕션 우위는 아닙니다.

Constraint-driven AI compiler scheduling R&D. 제약 기반 AI 컴파일러 스케줄링 R&D.

Result: Matched Exact DP's optimum on the benchmark — but Exact DP already solves it in ~7–10ms , so the WFC speed difference has no practical value. 벤치마크에서 Exact DP 최적값과 일치 — 하지만 Exact DP도 이미 약 7–10ms 에 풀어 WFC의 속도 차이는 실질적 가치가 없습니다.

The useful boundary was the cost model: its average correlation with measured GPU timing was only Spearman ρ=+0.52 — directional, not reliable enough for production scheduling. The differentiating 12KB SRAM spec exists on no real GPU; with A100/H100-scale SRAM the benchmark becomes trivial. Better calibration would require diverse hardware profiling data beyond this software-only experiment. Full result → 유의미한 경계는 cost model이었습니다. 실제 GPU timing과의 평균 상관은 Spearman ρ=+0.52 로, 방향성은 있지만 프로덕션 스케줄링 판단에는 부족했습니다. 차별화를 만든 12KB SRAM 스펙은 실제 GPU에 없고, A100/H100 규모 SRAM에선 문제가 trivial해집니다. 더 나은 교정에는 이 소프트웨어 실험 범위를 넘는 다양한 하드웨어 profiling 데이터가 필요합니다. 전체 결과 →