# Research Proposal Preparation Guide
## Surface Science Data Analysis & AI/ML Applications at TU Ilmenau

**Prepared for:** 3-month Research Internship Application  
**Group:** Technical Physics I, Prof. Stefan Krischok  
**Date:** March 2026

---

## Part 1: Understanding the Research Group (Explained Simply)

### What Does Prof. Krischok's Group Study?

Imagine you're a detective trying to figure out exactly what's happening on the surface of a material—not deep inside it, just the top few atomic layers (thinner than a human hair by a million times!). That's what Prof. Krischok's **Surface Science** group does.

**Their main question:** "What atoms and molecules are sitting on the surface of these special materials, how are they arranged, and how do they interact?"

**Why does this matter?** The surface is where all the action happens in real devices:
- Solar panels absorb light at their surface
- Batteries charge/discharge through surface reactions
- Computer chips work because of carefully controlled surface layers
- Catalysts (used in cars and factories) only work because of their surface chemistry

**Their specialty:** They focus on **Group III Nitrides** materials—especially:
- **GaN** (Gallium Nitride) → used in LED lights and power electronics
- **AlN** (Aluminum Nitride) → used in high-frequency electronics
- **ScGaN** and **AlScN** → new "smart" materials that can generate electricity when you squeeze them (piezoelectric)

These materials are like the "next generation" semiconductors that could replace silicon in many applications.

---

## Part 2: The Experimental Techniques (What the Machines Do)

Prof. Krischok's lab has five main "detective tools" to investigate surfaces. Let me explain each one like you're 15 years old:

### 1. **XPS (X-ray Photoelectron Spectroscopy)**

**What it does:** Shines X-rays at a surface and measures the electrons that fly off.

**Simple analogy:** Imagine throwing tennis balls (X-rays) at people wearing name tags (atoms). When a ball hits someone, their name tag flies off with a specific speed depending on who they are. By measuring the speed of the flying name tags, you can figure out "Who was standing there?" (which element) and "What were they wearing?" (their chemical bonds).

**What scientists measure:**
- **Which elements** are on the surface (Carbon? Nitrogen? Oxygen? Metal?)
- **What chemical state** they're in (Is it pure metal or is it oxidized? Is nitrogen bonded to metal or forming N₂ gas?)

**Data output:** Graphs with **peaks** at different energies. Each peak represents a specific element in a specific bonding state.

---

### 2. **UPS (Ultraviolet Photoelectron Spectroscopy)**

**What it does:** Similar to XPS, but uses ultraviolet light instead of X-rays.

**Simple analogy:** Same tennis ball idea, but now you're using softer balls that only knock off the outermost, loosest name tags. This tells you about the "energy levels" at the very surface.

**What scientists measure:**
- The **work function** (how much energy you need to pull an electron off the surface—important for batteries and solar cells)
- The **valence band edge** (the energy state of the outermost electrons—critical for understanding how electricity flows)

**Data output:** Graphs showing electron energy distribution, with specific features like the **secondary electron cutoff** and **valence band maximum**.

---

### 3. **LEED (Low-Energy Electron Diffraction)**

**What it does:** Shoots electrons at the surface and watches how they bounce back in patterns.

**Simple analogy:** Imagine bouncing a basketball on a tiled floor. If the tiles are perfectly arranged in a grid, the ball bounces in predictable patterns. If some tiles are missing or crooked, the pattern changes. LEED does this with electrons to "see" the atomic arrangement on the surface.

**What scientists measure:**
- **How atoms are arranged** on the surface (square grid? hexagonal? rotated?)
- Whether the surface is **clean and ordered** or **messy and contaminated**
- How **adsorbed molecules** (like water or oxygen) arrange themselves on the surface

**Data output:** **Images** with bright spots on a screen. The pattern of spots tells you the symmetry and spacing of surface atoms.

---

### 4. **TPD (Temperature-Programmed Desorption)**

**What it does:** Heats up a surface gradually and measures what molecules fly off at each temperature.

**Simple analogy:** Imagine you painted your room, and different things stuck to the walls: stickers, tape, glue, magnets. If you slowly heat the walls, the stickers peel off first (low heat), then tape, then glue (high heat), and magnets last (very high heat). By seeing what falls off at which temperature, you learn "what was stuck there and how strongly."

**What scientists measure:**
- **Which molecules** were sitting on the surface (water? oxygen? nitrogen?)
- **How strongly** they were bonded to the surface (binding energy)
- **How much** of each molecule was there (coverage)

**Data output:** **Time-series curves** showing gas intensity vs. temperature. Each peak in the curve represents a different molecule or binding site.

---

### 5. **IRAS (Infrared Reflection Absorption Spectroscopy)**

**What it does:** Bounces infrared light off the surface at a low angle and measures which wavelengths get absorbed.

**Simple analogy:** Different molecules "vibrate" in different ways (like different musical instruments make different sounds). Infrared light makes molecules vibrate. By seeing which colors of infrared light disappear (get absorbed), you can identify which molecules are on the surface and how they're oriented.

**What scientists measure:**
- **Chemical fingerprints** of surface molecules (CO? Water? Organic molecules?)
- **Molecular orientation** (Is the molecule standing up or lying flat?)
- **Surface reactions in real-time** (watch molecules react as you change temperature or add gases)

**Data output:** **Spectra** showing absorption peaks at specific wavelengths. Each peak is like a "fingerprint" for a specific molecular vibration.

---

## Part 3: The Real Data Analysis Problems (What Researchers Struggle With)

### Problem 1: Manual Peak Fitting Hell 🔥

**The current workflow for XPS/UPS data:**

1. Machine generates a raw spectrum file (often proprietary formats like `.sk`, `.spe`, `.vms`)
2. Researcher opens it in software like CasaXPS or Avantage
3. **Manually** draws a background curve under the peaks
4. **Manually** guesses where peaks should be and what shape they have
5. Clicks "fit" and tweaks parameters until it "looks good"
6. Repeats for 50-200 spectra from one experiment
7. Copies peak positions, areas, and widths into Excel spreadsheet

**Why this is painful:**
- **Time-consuming:** One sample with 5 elements can take 2-4 hours to analyze properly
- **Subjective:** Two researchers analyzing the same data will get different results depending on their "fitting strategy"
- **Error-prone:** Easy to miss overlapping peaks, use wrong peak shapes, or get trapped in local fitting minima
- **Not reproducible:** The original fitting parameters are rarely saved or shared in publications

**Real quote from expert:** *"There is a growing concern that the massive increase in XPS articles is accompanied by a **decrease in work quality** including meaningless chemical bond assignment."* [web:34][web:37]

**What goes wrong:**
- Wrong background subtraction (Shirley vs. Tougaard vs. linear—which one?)
- Ignoring **spin-orbit splitting** (elements like S, N, Cl produce doublet peaks that must maintain fixed intensity ratios and energy separations)
- Using **unrealistic peak widths** (all peaks from the same chemical state should have similar widths)
- Inventing too many chemical states to "force" the fit to support the paper's conclusion [web:37]

### Problem 2: The Hidden Metadata Crisis

**What actually happens in the lab:**

Sarah (PhD student) runs XPS on 15 GaN samples over 3 months:
- Sample names: `GaN_test1.sk`, `sample_march_v2.sk`, `final_FINAL_real.sk`
- Stored on her laptop in `Desktop/XPS_data/random_folder/`
- Critical info is in her paper notebook: "Sample prepared at 600°C, annealed 2h, measured March 15 at 10⁻⁹ mbar"

6 months later, her colleague Ahmed needs to compare his new AlScN samples to Sarah's old GaN data. He asks: *"Which sample was prepared at 600°C?"*

Sarah graduated. The knowledge is gone. The data is useless.

**The real problem:**
- **No standardized file naming**
- **No metadata embedded in files** (preparation conditions, measurement parameters, who ran it, when, why)
- **No shared database**—everyone has private folders on personal computers
- **No search capability**—can't filter by "all GaN samples measured between 500-700°C"
- **Data silos**—each student hoards their own data, no lab-wide sharing culture

### Problem 3: The Overlapping Peaks Nightmare

Many real samples have **multiple chemical states** overlapping in the same energy range:

**Example: Nitrogen 1s spectrum of a GaN surface exposed to air**
- N in GaN lattice: 397.0 eV
- N in oxidized surface: 398.2 eV  
- Adsorbed N₂ gas: 400.5 eV
- Contamination from lab air: 401.0 eV

All of these produce **Gaussian-shaped peaks** that blur together. The researcher must "deconvolve" them—but there are infinite ways to draw 4 overlapping Gaussians that fit the data equally well!

**Current solution:** Use chemical knowledge + comparison to reference spectra + trial-and-error.

**Problem:** This requires expertise that takes years to develop. And even experts disagree.

### Problem 4: LEED Pattern Interpretation

**Current workflow:**
1. Take photograph of LEED diffraction pattern on phosphor screen
2. **Manually** measure distances between spots with a ruler (or digital equivalent)
3. **Manually** calculate lattice parameters
4. **Manually** compare to reference patterns from textbooks

**Problems:**
- Spots may be faint, blurry, or have low contrast
- Some spots represent substrate, others represent adsorbate overlayer—researcher must mentally "separate" them
- Pattern changes with electron energy—researcher must take dozens of images at different energies and compare them all
- Human error in spot identification and distance measurement

### Problem 5: TPD Curve Interpretation

**Current workflow:**
1. Collect mass spectrometer data as temperature ramps up (typically 100-200 data points)
2. See multiple overlapping peaks in the curve (different binding sites or molecules desorbing at different temperatures)
3. **Manually** try to "deconvolve" the curve into multiple desorption events
4. Use Redhead equation or Polanyi-Wigner equation to extract binding energies

**Problems:**
- Overlapping peaks from different molecules or binding sites
- Baseline drift in mass spectrometer signal
- Kinetic parameters (reaction order, pre-exponential factor) are uncertain
- Difficult to automate—requires physicist's intuition

---

## Part 4: Where AI/ML Can ACTUALLY Help (Realistic Solutions)

### ⚠️ CRITICAL REALITY CHECK

**Before proposing "AI solutions," understand this:**

Most surface science labs **do NOT need**:
- ❌ Deep neural networks trained on ImageNet
- ❌ Transformer models for "semantic understanding"
- ❌ Reinforcement learning to "optimize experiments"
- ❌ Blockchain for "secure data sharing"

**What they ACTUALLY need:**
- ✅ Boring but essential **data engineering** (file parsing, metadata extraction, database design)
- ✅ Classical **signal processing** (background subtraction, smoothing, baseline correction)
- ✅ Simple **unsupervised learning** (PCA, NMF, clustering) to find patterns
- ✅ **Automated curve fitting** with proper constraints (not "AI magic," just good optimization)
- ✅ **Practical tools** that PhD students can actually use (Python scripts, not enterprise platforms)

---

## Part 5: Three Realistic AI/Data Science Projects for Your 3-Month Internship

### Project 1: **Automated XPS/UPS Peak Analysis Pipeline** ⭐ (HIGHEST IMPACT)

#### What You Would Build

A Python library that takes raw XPS/UPS spectra files and automatically:
1. **Reads proprietary file formats** (`.sk`, `.spe`, Avantage export formats)
2. **Extracts metadata** automatically (sample name, measurement date, photon energy, analyzer settings)
3. **Applies background subtraction** using multiple methods (Shirley, Tougaard, linear) with comparison
4. **Detects peaks automatically** using signal processing (derivative analysis, prominence thresholds)
5. **Fits peaks with physical constraints**:
   - Enforces spin-orbit splitting rules (fixed doublet ratios and separations)
   - Constrains peak widths to realistic ranges
   - Uses appropriate peak shapes (Gaussian-Lorentzian mix, asymmetric line shapes for metals)
6. **Quantifies uncertainty** using statistical bootstrapping
7. **Exports results** to standardized format (CSV + JSON metadata)

#### Why This Is Actually Needed

- Saves **100+ hours per PhD student per year**
- **Improves reproducibility** (same code always gives same result)
- **Reduces human bias** in peak fitting
- **Standardizes** analysis across the entire research group
- Directly addresses the "XPS data quality crisis" mentioned in recent literature [web:34][web:37]

#### Technical Approach (Not "AI Magic," Just Good Engineering)

**ML Component (20% of the work):**
- **Non-Negative Matrix Factorization (NMF)** to automatically separate overlapping peaks into components
- **Gaussian Mixture Models (GMM)** to find peak centers and widths
- **PCA** to identify outliers and anomalous spectra

**Classical Algorithms (80% of the work):**
- **Savitzky-Golay smoothing** for noise reduction
- **Baseline correction** using iterative polynomial fitting
- **Constrained non-linear least-squares** (Levenberg-Marquardt) for peak fitting with physical constraints
- **Chi-squared and R² statistics** for goodness-of-fit evaluation

#### Deliverables

- Python package `surfxps` (open-source, installable via `pip`)
- Command-line tool: `surfxps analyze sample.sk --element N --method auto`
- Jupyter notebook tutorials with real group data
- Documentation with "decision tree" for choosing analysis parameters

---

### Project 2: **FAIR-Compliant Laboratory Data Catalog** ⭐ (ADDRESSES DIRECT GROUP NEED)

#### What You Would Build

A lightweight, local **data management system** (NOT a complex database, NOT cloud-based) that:

1. **Watches** designated folders where instruments save data
2. **Automatically ingests** new files and extracts metadata:
   - Technique (XPS, UPS, LEED, TPD, IRAS)
   - Sample identifier
   - Measurement date/time
   - Instrument settings (photon energy, temperature, pressure)
   - Operator name
   - Project/experiment ID
3. **Stores** raw files + metadata in organized structure:
   ```
   /lab_data/
       /2026/
           /03_March/
               /XPS/
                   sample_GaN_001.sk
                   sample_GaN_001_metadata.json
   ```
4. **Provides simple web interface** (Flask/Streamlit app) where researchers can:
   - Search: "Show me all GaN samples measured at >600°C"
   - Filter by date, technique, sample type, operator
   - View thumbnail previews of spectra
   - Download data + metadata packages
   - See which samples are related (preparation batch, comparison studies)

5. **Follows FAIR principles**:
   - **F**indable: Searchable metadata, unique identifiers for each dataset
   - **A**ccessible: Clear file organization, standard formats
   - **I**nteroperable: JSON metadata, standard spectroscopy data formats
   - **R**eusable: Complete provenance (who, when, why, how)

#### Why This Is Actually Needed

Dr. Slimi explicitly said: *"We want to figure out how to share our group data so everyone has easy access. We have huge sets of data... filter data for us for easier data analysis."* [file:2]

This is a **direct, expressed pain point** from the group.

#### Technical Approach (Data Engineering, Not AI)

**Technologies:**
- **Python** for file parsing and metadata extraction
- **SQLite** for lightweight local database (no server setup needed)
- **Watchdog** library for monitoring directories
- **Flask or Streamlit** for simple web interface
- **HDF5 or NetCDF** for standardized spectroscopy data storage
- **DVC (Data Version Control)** for tracking data changes over time

**No AI/ML needed** for this project—it's pure data engineering and software design.

#### Deliverables

- Automated ingestion scripts running as background service
- Web dashboard accessible on lab network: `http://labserver:5000/data`
- User manual and training session for all group members
- Migration scripts to import legacy data from student laptops
- Integration with Project 1's output (automated analysis results stored in database)

---

### Project 3: **LEED Pattern Analysis & Spot Detection** ⭐ (COMPUTER VISION APPLICATION)

#### What You Would Build

A Python tool that analyzes LEED diffraction images to:

1. **Automatically detect diffraction spots**:
   - Use image processing to find bright spots on dark background
   - Handle varying brightness, noise, and camera artifacts
   - Distinguish real spots from noise or cosmic ray hits

2. **Measure lattice parameters**:
   - Calculate distances between spots
   - Determine symmetry (square, hexagonal, rectangular)
   - Identify substrate vs. adsorbate overlayer patterns

3. **Compare to reference database**:
   - Match observed patterns to known surface structures
   - Suggest possible interpretations (e.g., "consistent with (2×2) oxygen overlayer on GaN(0001)")

4. **Track pattern evolution**:
   - Analyze series of LEED images at different electron energies
   - Generate I-V curves (intensity vs. voltage) for quantitative analysis
   - Detect surface phase transitions or cleaning progress

#### Why This Is Actually Needed

- LEED is used daily to check surface quality before XPS/UPS measurements
- Manual interpretation is time-consuming and requires expertise
- Students often misidentify patterns, leading to incorrect conclusions about surface structure

#### Technical Approach (Classic CV + Minimal ML)

**Computer Vision (70%):**
- **Blob detection** using OpenCV (Laplacian of Gaussian, Difference of Gaussians)
- **Hough transform** for detecting geometric patterns
- **Fourier analysis** to extract lattice vectors
- **Peak finding** algorithms (scipy.signal.find_peaks with prominence filtering)

**ML Component (30%):**
- **Clustering** (DBSCAN) to group spots belonging to same lattice
- **Template matching** to compare against reference patterns
- **Outlier detection** to flag anomalous spots (contamination, surface disorder)

**No deep learning needed**—classical CV methods work excellently for high-contrast spot patterns.

#### Deliverables

- Python package `leed_analyzer`
- GUI application (PyQt or Tkinter) for interactive analysis
- Library of reference patterns for common GaN/AlN surfaces
- Batch processing script: `leed_analyzer process_folder --pattern_type hexagonal`
- Validation report comparing automated analysis to expert manual interpretation

---

## Part 6: How to Choose the Right Project

### Decision Matrix

| Criteria | Project 1 (XPS/UPS Pipeline) | Project 2 (Data Catalog) | Project 3 (LEED Analysis) |
|----------|------------------------------|-------------------------|---------------------------|
| **Addresses group's stated need** | ⭐⭐⭐ (peak analysis mentioned) | ⭐⭐⭐⭐⭐ (directly requested by Dr. Slimi) | ⭐⭐ (useful but not explicitly requested) |
| **Technical feasibility in 3 months** | ⭐⭐⭐ (doable, well-scoped) | ⭐⭐⭐⭐ (very achievable) | ⭐⭐⭐ (straightforward CV) |
| **Impact on group productivity** | ⭐⭐⭐⭐⭐ (saves massive time) | ⭐⭐⭐⭐ (enables collaboration) | ⭐⭐⭐ (nice to have) |
| **Uses your CS/DS skills** | ⭐⭐⭐⭐ (ML + signal processing) | ⭐⭐⭐⭐⭐ (data engineering) | ⭐⭐⭐⭐ (computer vision) |
| **Suitable for BSc student** | ⭐⭐⭐ (challenging but doable) | ⭐⭐⭐⭐ (very appropriate) | ⭐⭐⭐⭐ (good scope) |
| **Publishable results** | ⭐⭐⭐⭐ (journal paper potential) | ⭐⭐⭐ (software paper) | ⭐⭐⭐ (software paper) |

### Recommended Approach: **Hybrid Project 1 + Project 2**

**Title:** *"Automated Signal Processing and FAIR Data Centralization for Photoelectron Spectroscopy of Group III Nitrides"*

**Split your 3 months:**
- **Month 1:** Data Catalog (Project 2) — Build the infrastructure for ingestion, metadata, and sharing
- **Month 2:** XPS/UPS Pipeline (Project 1) — Core algorithms for background subtraction, peak detection, fitting
- **Month 3:** Integration + Documentation — Connect automated analysis to data catalog, write tutorials, train group

**Why this combination is optimal:**
- **Directly addresses** both of Dr. Slimi's stated needs (data sharing + analysis)
- **Builds on each other** (catalog makes analysis results findable and reusable)
- **Deliverable at each stage** (you can show progress monthly)
- **Realistic scope** for a BSc student in 3 months
- **High impact** on the entire research group

---

## Part 7: Key Success Factors

### What Will Make Your Proposal Stand Out

✅ **You listened:** You can quote Dr. Slimi saying "we want to share our group data" and "we have huge sets of data"

✅ **You did your homework:** You researched Prof. Krischok's actual work (ScGaN photoelectron spectroscopy, Group III nitrides)

✅ **You understand the real problem:** Not "AI for AI's sake," but solving actual workflow bottlenecks

✅ **You're realistic:** You're not promising AGI or magical solutions—you're offering practical engineering

✅ **You've scoped appropriately:** Your project is achievable in 3 months by a BSc student

✅ **You show mutual benefit:** You'll learn experimental physics + surface science, they get useful tools they'll actually use after you leave

✅ **You mention FAIR principles:** This is a huge buzzword in German academic physics right now (NFDI initiative)

✅ **You cite the literature:** Reference the "XPS data quality crisis" papers [web:34][web:37] to show you understand the field's challenges

### What to Avoid

❌ **Don't propose "deep learning for XPS"** without explaining why simpler methods won't work

❌ **Don't suggest cloud/enterprise solutions** (Databricks, Snowflake) for a small university lab

❌ **Don't promise to "revolutionize" their research**—be humble

❌ **Don't propose 10 different projects**—focus on ONE clear, achievable goal

❌ **Don't ignore physics constraints**—your ML model must respect spin-orbit splitting rules, etc.

❌ **Don't treat them as "clients"**—you're joining a research team, not consulting

---

## Part 8: Final One-Page Proposal Structure

Use this exact structure:

### 1. Title (1 line)
"Automated Data Processing and FAIR-Compliant Data Management for Surface Science of Group III Nitrides"

### 2. Abstract (3-4 sentences)
State the problem (scattered data, manual analysis), your solution (Python pipeline + data catalog), the approach (NMF/GMM for peaks, SQLite for metadata), and expected impact (time savings, reproducibility).

### 3. Background (2-3 sentences)
Prof. Krischok's group uses XPS/UPS/LEED/TPD/IRAS to study ScGaN and AlScN surfaces. Current workflow involves manual peak fitting and scattered data storage, limiting throughput and reproducibility.

### 4. Objectives (2 bullet points)
- **Objective 1:** Automated XPS/UPS signal processing with NMF-based peak deconvolution
- **Objective 2:** FAIR-compliant data catalog for shared, searchable access to experimental datasets

### 5. Methodology (3 bullet points)
- **Data Engineering:** Python file parsers, SQLite metadata database, web dashboard
- **Signal Processing:** NMF, GMM, constrained non-linear fitting with physical constraints
- **Validation:** Compare automated analysis to expert manual results on benchmark datasets

### 6. Timeline (simple table)
| Month | Focus | Key Deliverables |
|-------|-------|------------------|
| 1 | Data catalog infrastructure | Automated ingestion, metadata extraction, web interface |
| 2 | XPS/UPS analysis algorithms | Background subtraction, peak fitting, uncertainty quantification |
| 3 | Integration & documentation | Combined system, user manual, training workshop |

### 7. Expected Outcomes (3 bullet points)
- Open-source Python library for surface spectroscopy analysis
- Searchable laboratory data catalog following FAIR principles
- 50-80% reduction in manual analysis time for typical XPS datasets

---

## Part 9: References to Cite in Your Proposal

[1] Biesinger, M. C. (2022). "A step-by-step guide to perform x-ray photoelectron spectroscopy." *Journal of Applied Physics*, 132(1).  
→ Establishes best practices for XPS analysis that your tool would follow

[2] Major, G. H., et al. (2020). "Perspective on improving the quality of surface and material data analysis in the scientific literature with a focus on x-ray photoelectron spectroscopy (XPS)." *Journal of Vacuum Science & Technology A*, 38(6).  
→ Documents the "data quality crisis" your project addresses

[3] Wilkinson, M. D., et al. (2016). "The FAIR Guiding Principles for scientific data management and stewardship." *Scientific Data*, 3(1).  
→ Justifies your data catalog design choices

These three references show you understand the field's current challenges and how your project fits into broader initiatives for improving scientific data quality and sharing.

---

## Final Advice

**Remember:** Prof. Krischok and Dr. Slimi are physicists, not computer scientists. They don't need you to impress them with jargon. They need you to:

1. **Understand their actual problems** (you do—you listened in the call)
2. **Propose realistic solutions** (you are—practical Python tools, not AI hype)
3. **Show you can deliver** (you can—your projects are well-scoped for 3 months)
4. **Demonstrate enthusiasm** (you have—you've done your homework)

Your **biggest competitive advantage** over Master's candidates is that you're **NOT proposing a complex physics thesis**. You're offering something they desperately need: someone who can **build the data infrastructure and automation tools** that will make everyone else's research faster and better.

That's exactly what Dr. Slimi asked for on the phone: *"share our group data," "huge sets of data," "filter data for us for easier data analysis."* [file:2]

You heard him. Now show him you understand and can deliver.

**Good luck! 🚀**

---

*Document prepared: March 2026*  
*Based on: Call notes with Dr. Slimi, Prof. Krischok's publications, current XPS/UPS data analysis challenges in surface science literature*
