AI / Agent Engineering

Astronomy intelligence
you can audit

Five projects, one thread: bring the evidence-first discipline of physics into machine learning. A locally fine-tuned 8B astronomy LLM, a mobile GUI agent at ZTE (planner + worker A2A), a knowledge-graph & experiment-design agent with Peking University, an auditable research agent, and a lightweight survey classifier — each built to refuse claims the evidence doesn't support.

Dual-track RAG · audit click to advance

maoAstro — an 8B astronomy LLM, fine-tuned on one 12 GB GPU

An end-to-end domain LLM: 8 GB of ADS papers distilled into a two-stage fine-tune with my own architecture tweak — QLoRA + AttnRes — then wrapped in a four-step dual-track RAG. The whole pipeline (training and 4-bit serving) runs on a single RTX 3080 Ti 12 G, and it beats GPT-4o and AstroSage-8B on domain accuracy and hallucination rate.

Data layer

Collect knowledge

Top-journal literature — ApJ / MNRAS / A&A via ADS. ~8k papers · 8 GB.

  • Fill the domain-knowledge gap
→
Training layer

Two-stage fine-tune

Incremental pre-train (30k) → instruction tune (20k).

  • Core: QLoRA + AttnRes
→
Augment layer

Dual-track RAG

Hybrid recall (vector + BM25), bge rerank Top20→Top5.

  • Core: bge reranker
→
Serving layer

Lightweight web

vLLM, 4-bit AWQ, ~5 stable concurrent users.

  • Core: PagedAttention
Architecture innovation

AttnRes — rotating attention 90° into the depth dimension

Standard residual connections just add each layer's output to the last, with a fixed weight of 1. In a deep model, knowledge injected early fades before it reaches the top (the "deep-knowledge forgetting" that hurts long astronomy reasoning). AttnRes groups the 32 layers into 8 blocks and replaces that fixed accumulation with a learnable attention over depth — a block can read directly from any earlier block. Only tiny pseudo-query vectors are trainable (~10K params, +1–2 GB VRAM), zero-initialized so training never destabilizes. Toggle the mode and watch a knowledge signal survive to the top.

Three residual designs: (a) standard residuals add each block's output; (b) full-attention residuals; (c) block attention residuals, where an AttnRes operator α attends over earlier block states via a learnable pseudo-query w with keys/values from earlier blocks.
Fig. 1. The three residual designs. (a) standard residuals sum each block's output; (c) Block-AttnRes replaces that sum with an attention operator α whose query is a small learnable vector w and whose keys/values are the earlier block states.
Standard residual — fixed unit weight
hn = hn−1 + f(hn−1)

Each block adds only the previous block's output, so knowledge injected early is attenuated with depth — the deep-knowledge forgetting that hurts long reasoning.

AttnRes — learnable attention over depth
hn = ∑i<n αn,i hi,   αn = softmax(wn K⊤√d),   K = [h0 … hn−1]

A block reads directly from any earlier block. The query wn is a tiny pseudo-query — the only trainable part — zero-initialized so training starts exactly at the standard residual and never destabilizes.

Trainable footprint — QLoRA on the pseudo-query
W = W0 + BA,   B∈ℝd×r, A∈ℝr×k,  r ≪ d

Base weights W0 stay frozen & 4-bit (NF4); only the low-rank BA and the pseudo-queries update — ~10K trainable params, +1–2 GB VRAM.

Schematic of the Block-AttnRes idea: freeze the base weights, and LoRA-tune only the depth-attention pseudo-queries. On my astronomy benchmark this lifted domain-QA accuracy 86% → 89.2%, complex-reasoning 78% → 83%, and cut hallucination 9% → 6%.

Bar-and-line chart: astronomy-task accuracy (bars) versus relative VRAM (line) as the model is split into N = 16, 8, 4, 2 blocks. Accuracy is 90.4, 89.5, 86.2, 84 percent while relative VRAM drops 100, 50, 25, 12.5 percent.
Fig. 2. Block-count ablation — bars: astronomy-QA accuracy; line: relative VRAM of the depth-attention cache, as the 32 layers are grouped into N = 16 → 2 blocks.
Why 8 blocks

The sweet spot: near-peak accuracy at a quarter of the cache

Finer granularity (N=16) is marginally more accurate (90.4%) but pays full attention-cache cost. Coarser (N=2) collapses toward standard residual and drops to 84%. N = 8 keeps 89.5% while cutting the depth-attention memory to ~50% — which is exactly why the whole thing fits on a single 12 GB consumer card.

  • N=16 — 90.4% · 100% VRAM
  • N=8 — 89.5% · 50% VRAM  ← chosen
  • N=4 — 86.2% · 25% VRAM
  • N=2 — 84.0% · 12.5% VRAM
Data pipeline

8 GB → 30k clean corpus

LaTeX / reference stripping by regex, a BERT relevance classifier to keep astrophysics core, and MinHash dedup (128–2048 tokens). Instruction set: 20k samples (DS v3.1 + human check) mixing single-turn / multi-turn / reasoning / tool-calls.

Domain vocabulary

+2000 astronomy tokens

Extended the Qwen2.5-7B / astroMlab-8B tokenizer with 2000 domain terms. Proper-noun spelling errors dropped 18% → 3%; QA accuracy +5%.

Why QLoRA

≤10 G VRAM, low forgetting

Full fine-tune needs ≥40 G and forgets fast. Rank ablation: r=8 (attn only) 78%, r=64 (all linear) 86% — best, r=128 84%. Then AttnRes on top.

RAG that refuses

Hallucination 28% → 9%

Semantic chunking + fixed-length fill lifted recall 72% → 92%; a 4-step retrieve→rerank chain and fact constraints cut hallucination to 9%, and with no reference the model declines to answer.

Two-stage fine-tune
  1. ① Incremental pre-training30k unsupervised corpus · epoch 2 · lr 5e-5 Inject astronomy base knowledge into a Qwen2.5-7B / astroMlab-8B backbone.
  2. ② Instruction tune + AttnRes20k instructions · epoch 3 · lr 1e-4 · early-stop + dropout Instruction following & dialogue, with the depth-attention pseudo-queries learned on top.

Hardware: a single RTX 3080 Ti 12 G · Transformers + vLLM · Chroma + SQLite.

Dual-track RAG · 4-step chain
  1. 1Question → metadata filter (time / journal)
  2. 2Hybrid recall — vector + BM25 → Top 20 candidates
  3. 3Rerank — bge-reranker precision sort → Top 5
  4. 4Precise context → generate, refuse if unsupported

Semantic chunking (200–800 tok) + fixed fill lifted recall 72% → 92%; multi-hop via query decomposition + parent/child blocks 58% → 70%.

Evaluation

Benchmarked against GPT-4o and AstroSage-8B

60% automated (DS v3.1 scoring) + 40% blind human review by 3 astronomers over 200 samples. Click a metric to compare.

ModelDomain acc.HallucinationReasoningVRAM
maoAstro AttnRes89.2%6%83%12 G
GPT-4o82%18%76%no local
AstroSage-8B80.9%12%78%16 G

ZTE — a mobile GUI agent that operates the phone for you

At ZTE Corporation I build a phone-side GUI agent: you say what you want in natural language, and it drives the real UI — opening apps, tapping, typing, scrolling — to get it done. It is a genuine agent-to-agent (A2A) system with two specialised models: a Planner that understands your screen and installed apps and decomposes the task, and a self-trained 32B Worker that grounds each step into concrete touches. Orchestrated on the OpenClaw framework.

Planner agent

4B multimodal · plans

Reads a screenshot of the current screen and the set of installed apps, then a large LLM decomposes the request into an ordered chain of executable sub-goals — choosing which app and what to do next.

  • Multimodal screen + app-inventory understanding
  • High-level task → sub-goal chain
  • Re-plans when the screen doesn't match expectation
Worker agent

32B · operates self-trained

A GUI-grounding model I trained to convert one sub-goal + the current screen into a concrete action — tap, swipe, type — with pixel coordinates, then reports the resulting screen back to the Planner to close the loop.

  • Sub-goal + screenshot → grounded UI action
  • tap · swipe · long-press · type
  • Reports new screen → Planner verifies / advances

Screens are redrawn, not screenshots — but the loop is the real one: perceive → plan → ground → act → verify, with a human gate before it spends money.

Planner · 4B

Understands your installed apps

A compact 4B multimodal model perceives the live screen and the app inventory, so the plan is grounded in what's actually on your phone — not a generic script.

Worker · 32B

A GUI model I trained

The 32B worker is self-trained for UI grounding: sub-goal + screenshot → exact tap / swipe / type. Splitting perception-planning from execution is what makes it a real A2A system.

OpenClaw

The orchestration layer

Built on the OpenClaw agent framework — a typed message bus carries sub-goals down and screen states back up, with a verify-and-replan loop between the two agents.

On-device

Real UI, real control

The agent drives the actual interface end-to-end — cross-app tasks that chain several apps together, exactly the multi-step errands people don't want to do by hand.

Peking University — a knowledge-graph & experiment-design agent

As algorithm lead I built a materials-science research platform for battery-electrolyte R&D with Peking University. It ships as two frontends: a Lab Intelligent Agent System (a Graph-RAG assistant with Q&A, knowledge-graph and experiment-execution modes) and a Knowledge-Graph Visualization System (build, manage and quality-score the graphs). Below: the product framework, an experiment-design walkthrough, and the extraction pipeline behind it.

System 1 · :3001

Lab Intelligent Agent System

Graph Retrieval-Augmented System · 实验室智能体系统

问答 · Q&A 知识图谱 · KG 实验执行 · Execute
  • Answers grounded in a chosen graph + literature
  • Retrieval depth: quick · deep · wide-area reasoning
  • Experiment-design mode → protocol for the Nanocraft auto-synthesis platform
System 2 · :6700

Knowledge-Graph Visualization System

图谱可视化系统 · build · manage · assess

图谱可视化 图谱管理 质量评估
  • file / content / pipeline graph generation · meta-graph · merge
  • Sub-graph & shortest-path queries, JSON/CSV export
  • Quality score per triple: accuracy · triple-support · usefulness
实验室智能体系统Lab Agent electrolyte_2026 · 4 812 triples
User question · experiment-design mode
Design an experiment to determine whether FEC or VC more significantly improves SEI-film stability in a LiPF₆ / EC / DMC electrolyte.
Agent workflow
  1. ◇Analysis pathdecompose goal → sub-tasks
  2. ▤Graph-RAG retrievaldeep reasoning over KG + papers
  3. ⚗Generate protocolcontrols · cell build · testing
  4. ✓Validateevidence & consistency check
Knowledge graph · Leiden communities hover a node
Generated protocol → Nanocraft
Behind the graph

Schema-based extraction · a 3 + 1 pipeline with quality gates

Every triple in the graph is earned. A four-stage LLM pipeline extracts and then audits it, two dedup passes fold synonyms and PubChem CIDs together, and Leiden community detection plus LLM enrichment organize the result before it lands in Neo4j.

Stage 1
Entity recognition

Schema: entity_types, relation_types, attributes

Stage 2
Relation extraction

typed relations between entities

Stage 3
Attribute extraction

properties & measured values

Stage 4 · QC
LLM validation

score every triple: accuracy · triple-support · usefulness; drop hallucinations (θ=0.5)

Dedup 1
Union-Find

abbreviation synonym graph → lexicographically-smallest canonical name

Dedup 2
PubChem CID

local SQLite compound-ID alignment; drop redundant abbreviation triples

Community
hierarchical Leiden

graspologic partition → LLM names Top-20 members & Top-10 relations

Store
Neo4j import

MERGE nodes by (label,name), CREATE parallel edges, 500 triples / txn

The same architecture generalizes: literature → clean chunks → schema-based triples → audited, deduplicated, community-organized graph → Neo4j — reused across both the materials agent here and the astronomy Astro Agent below.

Astro Agent — an auditable, self-evolving research agent

My open-source flagship (github.com/wangnengdejiamao/Astro_Agent): a 16-node LangGraph that resolves a target, pulls multi-survey data, models it through three mandatory iterations, audits every claim, and — only when the QA gate clears — drafts an ApJ-style manuscript. It uses filesystem-as-memory for a diff-able audit trail, and a delegation loop that turns new methods from papers into tested, registered tools by handing the implementation to a second coding agent.

Acquire / plan Evidence gate Self-heal loop Draft Halt

Schematic of the real analysis_agent graph — node names mirror the implementation. The hard cap of three modelling iterations, two replans and two reflexion rewrites is exactly how the agent refuses to over-claim.

Self-evolving toolbox · delegation

A controlled auto-learning loop — it writes its own tools

When a run hits "I'm missing a tool", the orchestrator doesn't invent an answer. It packages the new method from a paper into a structured spec and delegates the implementation to Claude Code (claude_code_delegate) as a second coding agent. The tool is only registered as a new Skill when its tests pass — with a human review gate. Watch the hand-off.

Multi-stack toolbox

The technology toolbox

The agent is deliberately full-stack. Hover or tap a tool to see how it's used — grouped by the four architectural layers.

Application
Agent
Processing
Data
Tool detail

Hover a tool above to see how Astro Agent uses it.

Working memory

LangGraph AnalysisState

Session-level context and temporary results passed node-to-node.

Semantic memory

SQLite RAG (BM25)

Keyword search over the local white-dwarf literature — backs methodology decisions.

Structured memory

Two-layer KG

Entity graph + meta-graph over Neo4j / SQLite / JSON — supports method transfer and evidence chains.

Rule memory

Codex-style hard constraints

Every turn emits JSON — prevents behavioural drift and keeps output predictable.

Knowledge graph + dual-track RAG

The agent builds a typed knowledge graph — 8 entity types (paper · source · WD class · survey · instrument · method · model · parameter) linked by typed relations. Shown here is the graph of my own three papers; retrieval combines BM25/vector recall with multi-hop graph traversal — pick a query (or click a node) to watch it land on an entity and pull in its connected context.

ZTF light-curve classifier — a lightweight CNN

Where it started: turn every phase-folded ZTF light curve into a 224×224 image, then let a mobile-grade convolutional net read its shape. It sorts each star into EA (detached), EW (contact) or Non-EB — reaching 99.4–99.6% on 1,884 held-out curves. Pick a class and watch the forward pass; add noise and the softmax loses its nerve. Code: ztf_CNN_maoclassification.

Real pipeline: ZTF g/r photometry from IRSA, a Lomb–Scargle period, phase-folding, then a 1×224×224 render fed to GhostNet or MobileNetV2. Accuracies and the confusion matrix are the repository's reported values; the forward-pass animation is schematic.