← Back to Blog

ETO Scout: 5,953 Chinese tech articles summarised in English, free, and deliberately not a dataset

· 14 min read ETO Scout: 5,953 Chinese tech articles summarised in English, free, and deliberately not a dataset

There are 5,058,555 Chinese-language works in OpenAlex, out of 323,964,135 total. Just over one and a half percent. Whatever is being written, argued and decided about technology inside China, most of it is not reaching English-language readers through the usual academic plumbing.

ETO Scout is a small, narrow, free tool aimed straight at that gap. It reads a specific list of Chinese-language technology sources, has editors screen the articles, generates English summaries with machine learning, has humans check those summaries, tags them, and puts the result on a page you can filter. It will also email you the new ones.

When I loaded it on 9 August 2026 it held 5,953 articles.

That is a small number, and the smallness is on purpose. What makes Scout worth writing about is not the scale. It is that the people who built it published an unusually blunt list of things you should not do with it, and that list is more useful than most tools' entire documentation.

What it actually is

Scout describes itself as "ETO's discovery tool for Chinese-language writing on science and technology." The version on the app itself is slightly different and slightly more precise: "a discovery tool for Chinese-language news and commentary on technology issues."

Note what is missing from both. It is not a research paper index. It is not a translation service. It is not a database. It indexes news, commentary, op-eds, interviews and think tank output, and its job is to help an English speaker notice that something was said.

ETO is the Emerging Technology Observatory, a project of Georgetown University's Center for Security and Emerging Technology. CSET describes itself as "a policy research organization within Georgetown University's Walsh School of Foreign Service." Scout is maintained by CSET's data and analysis teams. Its documentation records an initial release on 20 September 2023.

What it costs

Nothing, and there is no account.

I did not create a login, did not supply an email, and did not hit a paywall or a metered limit. The web interface is fully open, the filters work, every article page has a stable shareable permalink, and the links out to the original sources are free.

The email digests are also free and need only an address. You can hold multiple subscriptions on the same address with different settings, which is a nicer piece of design than it sounds: you can take a daily feed of editors' picks and a separate monthly digest of everything tagged semiconductors, without one drowning the other. Digests go out daily on weekdays, weekly, or monthly, and if nothing new matches your filters, you simply do not get an email that period.

There is no paid tier. There is a donate link, which goes to CSET.

On who ultimately pays, the answer is philanthropy, at a scale that reframes the word "free." ETO's launch announcement says it directly:

We gratefully acknowledge funding support from the Musk Foundation and Open Philanthropy for these resources and ETO's broader platform.

CSET itself, the parent, was created by a single grant. Open Philanthropy, which has since rebranded as Coefficient Giving, recorded a grant of $55,000,000 over five years to Georgetown University, awarded January 2019, "to launch the Center for Security and Emerging Technology (CSET), a new think tank dedicated to policy analysis at the intersection of national and international security and emerging technologies." The grant record notes CSET was then led by Jason Matheny, the former IARPA director.

So Scout is free in the sense that somebody else already paid. It is a small, carefully made artifact sitting on top of one of the largest philanthropic bets ever placed on technology policy research.

And unlike a lot of think tanks, CSET does publish who pays. Not on its about page, where you would look first, but on its donate page, banded by size. Open Philanthropy sits alone in the $10,000,000 and above band. The Musk Foundation is in the $1,000,000 to $9,999,999 band alongside Google.org, the William and Flora Hewlett Foundation, the Patrick J. McGovern Foundation, NobleReach and Craig Falls. Further down sit the Alfred P. Sloan Foundation, the Rockefeller Foundation, the National Science Foundation, Apple, Microsoft, Nvidia, the Chan Zuckerberg Initiative, Scale AI, Schmidt Sciences and Leidos. ETO publishes no separate funder list of its own, so for ETO specifically the launch acknowledgement is the only direct statement.

The sources, named

This is the part that decides whether Scout is useful to you, and to their credit ETO lists every source with a description rather than saying "leading Chinese outlets."

As documented, Scout currently covers selected articles from:

  • 36kr (36氪): a Beijing-based, NASDAQ-listed platform aggregating technology content, aimed at Chinese companies, institutional investors and local governments
  • AI Tech Talk (AI科技评论): in-depth reporting on the AI industry and academia
  • Bulletin of the Chinese Academy of Sciences (中国科学院院刊): a CAS flagship journal on strategic issues in China's scientific development
  • CAICT (中国信息通信研究院): white papers and research from the think tank under China's Ministry of Industry and Information Technology
  • Caijing (财经): a widely read nonstate business, finance and technology outlet
  • CASIA Research Trends (中国科学院自动化研究所科研动态): updates from the Institute of Automation, a flagship CAS AI research institute
  • CSET curation: selections from CSET analysts' own current Chinese-language reading
  • EE Times China (电子工程专辑): electronics industry news, established 1993, owned by AspenCore
  • EEFocus (与非网): an electronics industry portal owned by Supplyframe, a Siemens subsidiary
  • Leiphone (雷锋网): a Shenzhen-based science and technology portal
  • S&T Daily (科技日报): a weekday newspaper published under the PRC Ministry of Science and Technology
  • SASTIND (国家国防科技工业局): the public website of China's State Administration of Science, Technology and Industry for National Defense
  • Zhidx (智东西): news and commentary on AI, consumer technology and the tech industry

Thirteen sources, twelve outlets plus CSET's own analyst reading, deliberately spanning state media, ministry think tanks, academy journals, and privately run commercial tech outlets. That spread is the editorial thesis: you are not reading one voice, you are reading a set of them that disagree with each other.

How an article gets in

The pipeline is documented step by step, which is rare enough to be worth repeating.

Editors screen new articles from those sources as they publish. The screening is explicitly subjective and articles get dropped for stated reasons: it turns out to be a video or an advert rather than an article; it summarises foreign developments without China-relevant commentary; it is a generic explainer; or the reviewing editor judges it would not meaningfully advance an English-language reader's understanding, because it is vague, redundant or padded.

Surviving articles then go through automated extraction and translation of basic metadata, title, author, source and date. A machine learning model drafts a short summary and a long one. Human editors check those drafts against the original articles and revise. The same editors apply technology, application and theme tags. Then a second editor reviews everything before publication.

So: machine drafts, two humans, published. Not machine-generated content with a disclaimer, and not a purely manual digest either. The web interface is updated several times a week.

Editors also flag a small number of items as editors' picks, described as informal, subjective judgements about what is most interesting, particularly things unlikely to filter through to English-language media. The documentation is careful to say picks do not imply endorsement of the article or the source.

Using it

The interface is one page. Reverse chronological list, filters down the left side.

Filters are dates, three tag families (technology, application, theme), an editors' picks toggle, and free-text search. Articles start collapsed showing title, source, date, tags and a one-line summary. Click one and you get the longer summary, extra Chinese and English metadata, and the outbound links.

Those outbound links are three: the original source, Google Translate, and the Wayback Machine. That third one is a thoughtful touch, because the original is often on a site that will not load well or at all from outside China, and sometimes will not load later at all.

A concrete taste of what turns up. From the front page on the day I looked, filtered to nothing at all:

  • An S&T Daily piece on Zhang Yonghe's two-decade effort behind the Tianguan satellites, using lobster-eye X-ray optics for time-domain astronomy
  • An EEFocus commentary arguing that after China's R&D spending caught up with the US on a purchasing-power basis in 2024, semiconductor competitiveness now turns on research allocation and commercialisation rather than money
  • An EEFocus article promoting AR glasses for Chinese policing, which Scout's own summary flags as touting "forward-looking capabilities rather than confirmed deployments"

That third summary is the tell. An editor read an industry puff piece and wrote a summary that tells you it is a puff piece. That is an editorial product, not a feed.

ETO's own suggested starting points are similarly practical: last month's articles on Chinese innovation in genomics, articles about technological indigenisation and decoupling, and editors' picks on semiconductors, AI and robotics.

To subscribe, you apply the filters you want and click through. The subscription dialogue arrives pre-populated with whatever you had filtered, which is the correct way to build that feature and almost nobody does it.

The list of things not to do

Here is the section I wish more tools shipped. ETO publishes, under the heading "What uses are not recommended," four things:

Do not draw conclusions about China as a whole. In their words, Scout "isn't comprehensive or representative of any one source or actor, let alone all Chinese-language writing on these topics." Thirteen sources, screened by editorial judgement, is a sample chosen for interest, not for representativeness. Anyone building a "what China thinks about AI" argument on Scout counts is doing something the tool's authors have told them not to do.

Do not cite the summaries. Cite the original articles. The summaries exist for skimming, are partly machine-generated, and may contain errors. ETO even supplies the citation format they would prefer: full citation to the original article, then "Located via the Emerging Technology Observatory Scout service at [link]."

Do not treat it as a substitute for reading the sources. There is no full text in Scout at all. Summaries will not capture most of the substance.

Do not do quantitative analysis. No bulk dataset of the articles or metadata is provided, and because the corpus is not comprehensive to begin with, counting things in it produces numbers that mean nothing.

I have read a great many tool documentation pages. Very few of them contain the sentence "we don't recommend using Scout for complex or quantitative analysis" about their own product.

The licence, and the one thing that will trip you up

Scout is covered by ETO's general terms of use, and those terms differ sharply from the other big open research resources.

ETO Resources are licensed CC BY-NC 4.0. Attribution required, non-commercial only. That is not CC0, and it is not CC BY. If you are building anything commercial, ETO material is off limits without specific permission.

And this clause, which is easy to miss and hard to work around:

ETO Resources may not be downloaded, copied, or extracted using web scrapers or other automated or semi-automated means.

There is no API, no bulk export, and no permission to make your own. Scout is a place you go and read, or a newsletter you receive. That is the whole product surface, on purpose.

If you have been building pipelines against open research infrastructure, this is a genuine adjustment. Harvard Dataverse defaults to CC0 and hands you a fully scriptable API. OpenAlex publishes its entire 324-million-record corpus under CC0 in a public S3 bucket. Scout is the opposite design: hand curated, human reviewed, small, and explicitly not machine readable.

Both designs are defensible. They are just answers to different questions.

The rest of ETO, which is where the data actually lives

If you arrived at Scout looking for something you can compute on, the answer is one of ETO's other eight tools. The full set of nine, all free:

  • AGORA: a living collection of AI-relevant laws, regulations, standards and governance documents from the US and worldwide
  • Country Activity Tracker: global research, patenting and investment in AI, as a dashboard
  • Chinese Technical Glossary: English translations and annotations for over 15,000 specialised terms from Chinese science and tech sources
  • Map of Patents: global intellectual property data
  • Map of Science: the world's research literature organised into nearly 92,000 clusters by citation and text similarity
  • PATHWISE: workforce and education metrics for emerging technology talent across US regions
  • Research Almanac: high-level trends in AI, biology, cybersecurity and semiconductor research
  • Supply Chain Explorer: the advanced chip supply chain
  • Scout: this one

Most of those sit on top of the Merged Academic Corpus, which ETO documents as containing "detailed information on over 280 million scholarly articles, combining data from public and private sources." The MAC draws on five sources: The Lens, which is commercial and closed, plus the open-access arXiv, Papers with Code, Semantic Scholar and OpenAlex. Because the Lens data is licensed from a commercial provider, the MAC is not publicly available in raw form. You interact with it only through the tools. Note the fifth name on that list. ETO's bibliometric backbone is partly built on the same open index the companion article in this set covers.

ETO's documentation of the MAC's limitations is worth reading on its own, and one of them bears directly on Scout's reason for existing: "The MAC's coverage of Chinese publications is incomplete. Although the MAC includes many Chinese publications, many others are only available in China-based journals that are not included in our data sources."

There is one exception to the no-bulk-data rule, and it is a good one. CSET publishes its classifier outputs, keyed to OpenAlex work identifiers, as a downloadable dataset on Zenodo, currently version 5.17.0 at about 2.8 GB, with fields flagging whether a work is AI, computer vision, NLP, robotics, cybersecurity, AI safety, chip design and fabrication, or large language model relevant. It is CC BY-NC 4.0, like everything else of theirs, and it updates monthly. If you want CSET's judgement applied at scale to the scholarly literature, that file is the way to get it.

Who this is actually for

Scout is for a person, not a pipeline.

If you are an analyst, a journalist, a policy staffer, or an engineer trying to keep half an eye on what Chinese industry and state media are saying about semiconductors or AI, and you do not read Chinese, this is a genuinely good use of ten minutes a week. Subscribe to editors' picks. Read the summaries. Click through on the two or three that matter and run the original through translation.

If you were hoping to count things, measure sentiment, chart trends, or feed a model, Scout is not that and its authors will tell you so in writing. Go to Map of Science, the Country Activity Tracker, or the Zenodo file.

And the honest caveat about freshness: the newest article in Scout when I checked on 9 August 2026 was dated 31 July. The documentation says the interface updates several times a week, so that gap may just be what a quiet fortnight looks like in a hand-screened tool. It is worth knowing before you treat it as a breaking-news source. It is not one, and it does not claim to be.

Fact-check notes and sources

Counts and dates were measured on 9 August 2026 and will have moved since.

  • 5,953 articles held, the newest dated 31 July 2026, and the filter and interface behaviour described: measured directly in the Scout web interface
  • Self-descriptions, the "skimming and discovery" and machine-translation cautions: scout.eto.tech and ETO documentation for Scout
  • Full source list with descriptions, the screening criteria, the summarisation and two-editor review pipeline, editors' picks, the update cadence and the email digest behaviour: ETO documentation for Scout
  • The four "uses not recommended" and the suggested citation format: same page, under "What uses are not recommended?" and "How do I cite it?"
  • Initial release 20 September 2023, and the credits naming Zach Arnold for concept and documentation and Neha Singh and Jennifer Melot for engineering and maintenance: same page, major change log and credits
  • CC BY-NC 4.0 licensing and the prohibition on scrapers and automated or semi-automated extraction: ETO Terms of Use
  • ETO as a project of CSET, and CSET as a policy research organization within Georgetown's Walsh School of Foreign Service: eto.tech and CSET About Us
  • Musk Foundation and Open Philanthropy funding acknowledgement: Introducing the Emerging Technology Observatory
  • $55,000,000 over five years, awarded January 2019, to launch CSET, then led by Jason Matheny: Open Philanthropy's grant record for Georgetown University, Center for Security and Emerging Technology. Open Philanthropy now operates as Coefficient Giving and the original grant URL redirects to a fund landing page, so the archived snapshot is the citable copy.
  • CSET's banded donor list, placing Open Philanthropy above $10,000,000 and the Musk Foundation between $1,000,000 and $9,999,999: CSET Donate, under "Thank you to our donors!"
  • The full nine-tool list and their descriptions: ETO Our Tools
  • Merged Academic Corpus at over 280 million articles, its five sources (The Lens, arXiv, Papers with Code, Semantic Scholar and OpenAlex), why it is not publicly available, and the stated limitation on Chinese publication coverage: ETO documentation for the Merged Academic Corpus
  • Map of Science at nearly 92,000 clusters: ETO documentation for the Map of Science
  • CSET metadata over OpenAlex works, version 5.17.0, about 2.8 GB, CC BY-NC 4.0, monthly updates, and the classifier field list: Zenodo record 20370823 and the cset_openalex repository
  • 5,058,555 Chinese-language works out of 323,964,135 total: measured via https://api.openalex.org/works?group_by=language

ETO does not publish its own funder list. CSET publishes a banded one on its donate page rather than its about page, and I have quoted the bands rather than inferring exact amounts. The observation about the newest article being nine days old on the day I checked is mine, not an ETO statement.

Related reading

This post is informational. Mentions of the Emerging Technology Observatory, CSET, Georgetown University and the named Chinese-language outlets are nominative fair use. No affiliation is implied.

← Back to Blog

Accessibility Options

Text Size
High Contrast
Reduce Motion
Reading Guide
Link Highlighting
Accessibility Statement

J.A. Watte is committed to ensuring digital accessibility for people with disabilities. This site conforms to WCAG 2.1 and 2.2 Level AA guidelines.

Measures Taken

  • Semantic HTML with proper heading hierarchy
  • ARIA labels and roles for interactive components
  • Color contrast ratios meeting WCAG AA (4.5:1)
  • Full keyboard navigation support
  • Skip navigation link
  • Visible focus indicators (3:1 contrast)
  • 44px minimum touch/click targets
  • Dark/light theme with system preference detection
  • Responsive design for all devices
  • Reduced motion support (CSS + toggle)
  • Text size customization (14px–20px)
  • Print stylesheet

Feedback

Contact: jwatte.com/contact

Full Accessibility StatementPrivacy Policy

Last updated: April 2026