Sunday, September 13, 2026

請幫我列舉,重要的數位人文學者及其主要作品,謝謝,(ChatGPT, 2026-9-13)

 可以。若把「數位人文」(Digital Humanities, DH)理解為:利用計算、資料庫、文本探勘、網絡分析、GIS、視覺化與數位典藏,重新提出人文學問題,那麼它其實不是一個單一學派,而是一大片彼此重疊的研究群。早期叫 humanities computing,大約 2000 年後才逐漸以 digital humanities 為通稱。

我會把重要人物分成六條線來看。

1. 「祖師爺」:Roberto Busa(1913–2011)

Busa 是義大利耶穌會神父,一般被視為 humanities computing 的奠基者之一。1940 年代開始與 IBM 合作,把 Thomas Aquinas 等人的拉丁文著作做成機器可檢索的 Index Thomisticus;整個計畫處理逾一千萬詞,歷時數十年。

重要作品/計畫:

  • Index Thomisticus
  • 其真正重要性不是「做了一個索引」,而是第一次大規模問:如果整個思想家的 corpus 都能計算,閱讀會變成什麼?

所以可以把他看成:

close reading → corpus → machine-readable humanities


2. 理論奠基者:Willard McCarty、John Unsworth、Jerome McGann

這幾位的重要性,在於把「電腦幫忙整理文本」提升成:計算本身是否可以成為人文學的認識論?

Willard McCarty 的代表作是:

  • Humanities Computing(2005)
  • “Tree, Turf, Centre, Archipelago—or Wild Acre?” 等文章

McCarty 最值得讀的不是技術,而是他不斷問:

modelling 到底是什麼?

模型不是把人文現象「忠實複製」,而是有意識地簡化、變形,再藉由模型失敗之處重新認識研究對象。

John Unsworth 的經典概念是:

  • “Scholarly Primitives”
  • “What Is Humanities Computing and What Is Not?”

他提出閱讀、搜尋、比較、註釋、引用、發現等「scholarly primitives」——也就是不論紙本或數位,人文學者其實反覆做的一些基本動作。Unsworth、McCarty、McGann 都已被收入 DH 的核心文獻譜系。

Jerome McGann 則由文本編輯、書籍史進入數位文本:

  • Radiant Textuality: Literature after the World Wide Web(2001)
  • Rossetti Archive

他最重要的觀念之一是:文本不是只有「文字內容」;版面、版本、物質形式、編輯史,也都是文本的一部分。

這一點其實非常重要:數位化不應該只是把書變成 .txt


3. 文學的大資料革命:Franco Moretti、Matthew Jockers、Ted Underwood

這是今日最容易辨認的 DH 路線。

Franco Moretti

代表作:

  • Graphs, Maps, Trees(2005)
  • Distant Reading(2013)

Moretti 創造了最有名的 DH 術語:

distant reading

不是讀十本小說,而是研究十萬本小說。

例如問:

  • 某種小說類型何時出現、衰亡?
  • 小說標題百年來如何變化?
  • 文學形式如何像物種一樣演化?
  • peripheral literature 如何進入 world literature?

Moretti 未必是最「技術型」的 DH 學者,卻可能是思想影響最大的一位,因為他改變了研究尺度。

Matthew L. Jockers

如果 Moretti 提出 distant reading,Jockers 就把它做成比較完整的方法論。

代表作:

  • Macroanalysis: Digital Methods and Literary History(2013)
  • Text Analysis with R for Students of Literature(2014)

他稱自己的方法:

macroanalysis

以數千、數萬本小說進行:

  • topic modeling
  • stylometry
  • authorship attribution
  • sentiment
  • literary influence

Stanford 早期就把他的研究描述為以數千文本追蹤跨世紀文學史與風格趨勢。

Ted Underwood

我特別推薦。

代表作:

  • Why Literary Periods Mattered(2013)
  • Distant Horizons: Digital Evidence and Literary Change(2019)

Underwood 比第一代 distant reading 更謹慎。他關心:

機器學習發現的 pattern,到底能不能成為「歷史證據」?

他用 classification、regression、machine learning 研究小說類型、性別、文學聲望與歷史變化。

如果 Moretti 是 manifesto,Jockers 是 method,Underwood 很像是 epistemology after the method


4. 圖像、界面與文化大資料:Lev Manovich、Johanna Drucker

這條線可能尤其值得您注意,因為它已經不只是「讀文字」。

Lev Manovich

代表作:

  • The Language of New Media(2001)
  • Software Takes Command(2013)
  • Cultural Analytics(2020)
  • AI Aesthetics(2018)
  • Artificial Aesthetics(2024)

Manovich 提出的核心詞:

cultural analytics

問題不是「如何讀十萬本書」,而是:

How can we see a billion images?

他把數百萬張 Instagram 圖片、電影畫格、藝術品、視覺風格當成資料集,用 computer vision、statistics、visualization 來觀察文化。

這是 DH 與今日 AI 最直接相接的地方之一。


Johanna Drucker

如果要我選一位「最人文、最不會被計算迷惑」的 DH 理論家,我會選 Drucker。

代表作:

  • SpecLab: Digital Aesthetics and Projects in Speculative Computing(2009)
  • Digital_Humanities(與 Burdick、Presner、Lunenfeld、Schnapp 合著,2012)
  • Graphesis: Visual Forms of Knowledge Production(2014)
  • Visualization and Interpretation(2020)
  • The Digital Humanities Coursebook(2021)

她最有名的批判之一是:

人文資料不是 data,而更接近 capta

即不是「天然存在、等著被抓取的資料」,而是研究者依某個觀點取出來的東西

所以她反對把 visualization 當成透明的「真相顯示器」。

對她而言:

圖表也是論述。
database 也是 interpretation。
interface 也是 epistemology。

我以為這一點對 AI 時代會越來越重要。


5. 「電腦不是替代詮釋,而是刺激詮釋」:Geoffrey Rockwell & Stéfan Sinclair

代表作:

  • Hermeneutica: Computer-Assisted Interpretation in the Humanities(2016)
  • Voyant Tools

這兩位發展的 Voyant Tools,今日仍是最容易入門的文本分析工具之一。

但他們真正有意思的觀點是:

computer-assisted interpretation

不是:

電腦算完 → 得出答案。

而是:

人讀 → 電腦發現奇怪 pattern → 人回去重讀 → 改問題 → 再計算。

所以 computation 成為一種hermeneutic provocation。MIT Press 對此書的定位也是「computer-assisted interpretive practices」。

這其實比「AI 幫我摘要一本書」深很多。


6. DH 的自我批判:Alan Liu、Tara McPherson、Roopika Risam、Lauren Klein

2010 年後數位人文開始反省:

誰的資料?
誰建立 archive?
哪些人被排除?
algorithm 有沒有政治?
數位化是不是把舊殖民權力重新 digitize 一次?

Alan Liu

代表作:

  • The Laws of Cool(2004)
  • Friending the Past(2018)

Liu 特別關心 digital culture、知識工作與 humanities 的公共角色。

他問的不是「新科技好不好」,而是:

數位知識制度正在把「人文」變成什麼?


Tara McPherson

代表作:

  • Feminist in a Software Lab: Difference + Design(2018)
  • Reconstructing Dixie(2003)
  • Vectors / Scalar 數位出版計畫

她把 race、gender、software architecture、interface design 放在一起思考。

核心問題很尖銳:

我們能不能一邊宣稱 software 是中性的,一邊忽略它產生於什麼社會結構?


Roopika Risam

代表作:

  • New Digital Worlds: Postcolonial Digital Humanities in Theory, Praxis, and Pedagogy(2018)

她發展 postcolonial digital humanities

尤其重要的一點是:

殖民檔案本身已經有偏見;
把殖民檔案 digitize,並不會自動去殖民。

甚至可能只是:

digitizing colonialism

她因此討論 colonial violence、digital archives、NLP 與「human」概念如何被預設。


Lauren Klein(以及 Catherine D’Ignazio)

代表作:

  • Data Feminism(2020)
  • Klein 另有 An Archive of Taste(2020)

Data Feminism 雖跨 data studies/DH,但現在幾乎已是 critical DH 必讀。

核心不是「女性怎麼用資料」,而是:

data science 本身如何牽涉權力?

例如:

  • 哪些東西被計數?
  • 哪些東西從未進入 database?
  • 誰有權分類?
  • missing data 本身說了什麼?


還有一群是「數位人文史與學科建制」必須知道的人

Susan Schreibman、Ray Siemens、John Unsworth 主編的:

A Companion to Digital Humanities(2004)

幾乎可視為 DH 正式成為一個學術領域的重要里程碑。它從 archaeology、linguistics、lexicography、literary studies 一路涵蓋到 textual analysis。

Melissa Terras、Julianne Nyhan、Edward Vanhoutte 主編:

Defining Digital Humanities: A Reader(2013)

則特別適合拿來看:

humanities computing 是怎麼慢慢變成 digital humanities 的。

編者自己就說,此書刻意收集 humanities computing 與 DH 的核心論爭,而不是硬給出一個唯一的定義。


如果只讀「十本」,我會這樣排

  1. Roberto Busa — Index Thomisticus(不是書,而是源頭)
  2. Willard McCarty — Humanities Computing
  3. Schreibman / Siemens / Unsworth eds. — A Companion to Digital Humanities
  4. Franco Moretti — Graphs, Maps, Trees
  5. Franco Moretti — Distant Reading
  6. Matthew Jockers — Macroanalysis
  7. Johanna Drucker — Graphesis
  8. Lev Manovich — Cultural Analytics
  9. Geoffrey Rockwell & Stéfan Sinclair — Hermeneutica
  10. Roopika Risam — New Digital Worlds

讀完這十部,差不多會看到 DH 的一整個運動:

索引 → corpus → model → distant reading → macroanalysis → visualization → cultural analytics → hermeneutics → critical / postcolonial DH

而我覺得眼下,到了 LLM 時代,數位人文正在發生第三次轉折

第一階段是:

digitization
把人文資料數位化。

第二階段是:

computation
在幾十萬、幾百萬文本之間尋找 pattern。

現在第三階段開始變成:

machine-assisted synthesis
讓 AI 在龐大而混亂的 archive 中,尋找尚未被作者本人辨認出來的關聯。

這一階段有趣之處,恰恰已經不是 Moretti 的 distant reading。我會暫時叫它 latent reading:不是站遠一點看整個 corpus,而是讓機器進入 corpus 的潛在關係空間,問:

這四萬六千篇碎片之間,有哪些關係,是作者二十二年來自己也不知道的?

從這個角度看,您先前所說的 blog archive,其實是一個非常典型、又比傳統 DH 更進一步的研究材料:不只是 digital archive,而是 longitudinal intellectual corpus。Moretti 研究一個文類的生成;這裡則可以研究「一個人的思想如何在二十二年間生成、消失、歧出、回返」。這大概已經站到 digital humanities 與 AI humanities 的交界處了。