Documentation
How the data is collected, what the rankings mean, and known limitations.
1. Data Sources & Joins
Each record combines three kinds of source: the university's official staff list, the university's own publication records supplemented by open indexes, and journal ranking lists.
Researcher profiles
For every university, the staff list comes from the official Accounting and Finance staff directory. Each researcher's profile provides their name, job title (from which the academic level is derived), field of research, profile URL and, where the university publishes them, an ORCID iD and a repository ID (source_id). Staff with no publications are left out.
Publications
Publications come first from each university's own system, which is treated as the official record:
- UQ — UQ eSpace, looked up by each person's eSpace author ID.
- UWA — UWA's Pure research portal: each person's publication feed and publication pages.
- Monash — Monash's Pure research portal: each person's publication feed. OpenAlex then adds DOIs and ISSNs to the portal's papers by exact title match, using only OpenAlex records whose ORCID matches the researcher.
- Melbourne — Minerva Access, the University of Melbourne repository.
- ANU, UNSW and Sydney — each researcher's official staff profile page.
- Adelaide — the journal articles listed on each researcher's Adelaide research profile, each with its DOI.
Three open indexes then add papers the university's records miss: ORCID (the researcher's own public record), Crossref and OpenAlex. They are looked up by ORCID iD, and the OpenAlex search is restricted to papers written while at that university. Duplicate copies of the same paper are merged by DOI, or by title and year where there is no DOI. Only journal articles are kept.
Co-authored papers are counted once per matching author, so a paper with two staff co-authors appears twice in the database.
Journal rankings
Journal quality data comes from three sources, matched to publications by ISSN (the ABDC list falls back to an exact title match when a paper has no ISSN):
- ABDC Journal Quality List 2025 — from
data/ABDC-JQL-2025-v1-260326.xlsx. Providesquality_rank(A*, A, B, C) andabdc_edition. - Scimago Journal Rankings 2025 — from
data/scimagojr 2025.csv. Providessjr,sjr_quartile,h_indexandcites_per_doc_2y. - Clarivate Journal Impact Factor (JCR) — fetched from the Clarivate Journal Citation Reports API when an API key is configured. Provides
impact_factor,impact_factor_5yrandjcr_year. Only journals indexed in Web of Science have an impact factor; for the rest these are null.
2. Standardisation Pipeline
One command runs the whole pipeline for a university: python run.py --uni <name>, or python run.py --all for all eight. Each university goes through the same steps:
Step 1 — Collect
The university's adapter (base_scrapers/<name>.py) reads the official staff directory and the university's own publication records. Reviewed identity corrections in data/ (for example a verified ORCID) are applied here.
Step 2 — Supplement
ORCID, Crossref and OpenAlex are searched for papers the university's records miss.
Step 3 — Clean and enrich
Titles, years, DOIs and publication types are cleaned; errata, editorials and reviewed exclusions (data/publication_exclusions.csv) are removed. OpenAlex and Crossref then add citation metrics, open-access status and missing ISSNs by DOI.
Step 4 — Rate journals
Each journal is matched to the ABDC list, Clarivate Journal Citation Reports (when an API key is configured) and Scimago.
Step 5 — Screen
screen.py removes open-index papers that probably belong to someone else with the same name: when a researcher has at least five open-index journal articles and fewer than a quarter of them are in ABDC-rated journals, those papers are set aside. Some specific cases are also handled by hand-reviewed rules. Papers from the university's own records are never removed. Everything set aside is listed in <uni>_screened_out.csv with the reason.
Step 6 — Export
The results are written to final output/<uni>/ as four tables — staff, publications, journals and harvest — each as CSV and JSON.
Step 7 — Load
load.py rebuilds the website's SQLite database (site/research.db) from every university's folder in final output/. The admin page's Refresh button runs the pipeline for one or all universities and then this step, replacing the live database only if everything succeeds.
OpenAlex percentile
Where available, OpenAlex returns a citation_normalized_percentile value indicating a publication's citation standing among comparable works. This is stored as citation_percentile in the database.
3. Database Schema
The database (site/research.db) is a SQLite file managed by SQLAlchemy and rebuilt by load.py. It has three tables in use, below. A fourth, harvest, is defined but not currently filled; when each source was last harvested is recorded in each university's <uni>_harvest file instead.
Researcher
| Column | Type | Notes |
|---|---|---|
researcher_id | INTEGER PK | Auto-increment |
name | TEXT | Name with titles removed |
job_title | TEXT | Position as the university lists it; the staff CSV adds academic_title (from the level) and admin_title (e.g. "Dean") |
academic_level | TEXT | A–E, derived from the job title; null when it doesn't map |
university | TEXT | University name (see §8) |
field_of_research | TEXT | Accounting or Finance |
source_id | TEXT | ID in the university's own system, where published |
orcid | TEXT | ORCID iD, where available |
profile_url | TEXT | Link to the official university profile |
Journal
| Column | Type | Notes |
|---|---|---|
journal_id | INTEGER PK | |
journal_name | TEXT | ABDC title where listed, otherwise the name as found; unique |
journal_raw | TEXT | Name exactly as the source gave it |
publisher | TEXT | |
issn | TEXT | One or more ISSNs, separated by ; |
quality_rank | TEXT | ABDC: A*, A, B or C; null if not on the list (see §4) |
abdc_edition | TEXT | E.g. "2025" |
impact_factor | REAL | Clarivate 2-year JIF; null outside Web of Science |
impact_factor_5yr | REAL | Clarivate 5-year JIF |
jcr_year | INTEGER | JCR release the impact factors come from |
sjr | REAL | Scimago Journal Rank score |
sjr_quartile | TEXT | Q1–Q4 per Scimago |
h_index | INTEGER | Journal h-index per Scimago |
cites_per_doc_2y | REAL | Scimago 2-year citations per document |
scimago_year | TEXT | Scimago release the figures come from |
Publication
| Column | Type | Notes |
|---|---|---|
publication_id | INTEGER PK | |
researcher_id | INTEGER FK | → researcher.researcher_id |
journal_id | INTEGER FK | → journal.journal_id (nullable) |
title | TEXT | |
year | INTEGER | |
author_count | INTEGER | Including external co-authors |
authors | TEXT | All authors, separated by ; |
doi | TEXT | |
article_url | TEXT | DOI link, or the source's link when there is no DOI |
link | TEXT | Link given by the source |
quality_rank | TEXT | The journal's ABDC rating, copied for convenience |
sjr_quartile | TEXT | The journal's Scimago quartile, copied for convenience |
citation_percentile | REAL | OpenAlex citation percentile, 0–1 |
cited_by_count | INTEGER | From OpenAlex |
fwci | REAL | OpenAlex Field-Weighted Citation Impact |
oa_status | TEXT | gold, diamond, hybrid, bronze, green or closed |
oa_url | TEXT | Link to a free copy, where one exists |
publication_status | TEXT | published or forthcoming (UWA rows use UWA's own wording, e.g. "Published - Jun 2024") |
source | TEXT | The university system or open index the record came from |
4. Ranking Metrics
ABDC Journal Quality List
The Australian Business Deans Council (ABDC) Journal Quality List rates journals in business and economics. The 2025 edition is used. The quality_rank field takes one of four values, or is blank when the journal is not on the list — common and expected for journals outside business, such as medical or engineering journals:
Scimago Journal Rankings (SJR)
Scimago provides three metrics sourced from Elsevier Scopus data:
sjr— weighted citation score; accounts for the prestige of the citing journals.sjr_quartile— Q1 (top 25%) to Q4 (bottom 25%) within the journal's subject category.cites_per_doc_2y— average citations per document over the preceding two years (analogous to a 2-year impact factor).
Clarivate Journal Impact Factor (JCR)
impact_factor is the Clarivate 2-year Journal Impact Factor from Web of Science Journal Citation Reports, and impact_factor_5yr the 5-year version; jcr_year is the JCR release they come from. They are null for journals outside Web of Science. Scimago's cites_per_doc_2y is a similar measure with wider coverage.
OpenAlex Citation Percentile
citation_percentile is OpenAlex's citation percentile for the individual publication, not the journal, normalised by publication year, type and subfield. It is stored as a fraction from 0 to 1: a value of 0.95 means the paper is cited more than 95% of comparable papers.
5. Column Guide
A short guide to the columns in the data. A blank value always means not available, never zero: a journal with no ABDC rating is simply not on the ABDC list.
Researchers
| Column | Meaning |
|---|---|
academic_level | Australian academic level from the job title: A Associate Lecturer, B Lecturer, C Senior Lecturer, D Associate Professor, E Professor. |
field_of_research | Accounting or Finance, from the department the university lists the person under. |
orcid | A permanent researcher identifier, used to find their publications reliably. |
Publications
| Column | Meaning |
|---|---|
cited_by_count | How many times the article has been cited. |
fwci | Field-Weighted Citation Impact: citations compared with similar articles of the same field, year and type. 1.0 is the world average; 2.0 is twice as many citations as expected. |
citation_percentile | Where the article ranks by citations among comparable articles, from 0 to 1. 0.91 means cited more than about 91% of them. |
oa_status | Whether it is free to read: gold, diamond, hybrid or bronze (free from the publisher), green (a free copy in a repository), or closed. |
source | Where the record came from: the university's own system (the official record) or an open index (ORCID, Crossref, OpenAlex). |
Journals
| Column | Meaning |
|---|---|
quality_rank | ABDC rating: A* (highest), A, B or C. |
impact_factor | Clarivate Journal Impact Factor: average citations in a year to the journal's articles from the previous two years. Only journals in Web of Science have one. |
impact_factor_5yr | The same over five years; steadier, and better suited to slower-citing fields such as accounting. |
sjr | SCImago Journal Rank: citations weighted by the prestige of the citing journals. Higher is better. |
sjr_quartile | Q1 (top 25% of its subject category) to Q4 (bottom 25%). |
h_index | The journal's h-index: h of its articles have each been cited at least h times. |
The full data dictionary explains every column in every table, where each value comes from, and how to read the metrics.
Download the full data dictionaryA Markdown text file; it opens in any text editor.
Download the full dataset
Every university in one file per table. Each publication row carries the
researcher's name, university and ORCID and the journal's name and ISSN, so it can be
filtered or joined to another spreadsheet on its own. researcher_id
and journal_id link the three files together; ORCID is the
safest key for matching researchers to your own records, because names can collide
and vary in spelling.
CSV files; they open directly in Excel.
6. Known Limitations
<uni>_screened_out.csv, but a researcher whose genuine output is large can still carry some wrong papers. Corrections belong in the files in data/, not in the outputs, which the next run overwrites.
7. Worked Example
A researcher at Monash University publishes a paper in the Journal of Finance (ISSN 0022-1082). Here is how the record is assembled:
- Collect. The Monash adapter finds the researcher in the Accounting staff directory, reads their ORCID iD from their Monash research profile, and finds the paper in their profile's publication feed.
- Identifiers. OpenAlex is searched by the researcher's ORCID iD, using only records that carry that exact ORCID. The OpenAlex copy with the same title supplies the paper's DOI and the journal's ISSN,
0022-1082. - Enrich. OpenAlex adds the citation count, FWCI, citation percentile and open-access status, looked up by DOI.
- Rate the journal. ISSN
0022-1082is found on the ABDC list asA*, in Clarivate JCR (impact_factor,impact_factor_5yr) and in Scimago (sjr,sjr_quartile,h_index,cites_per_doc_2y). - Screen. The paper came from Monash's own records, so the screening step leaves it alone.
- Export and load. It is written to
final output/monash/monash_publications.csv, with the journal inmonash_journals.csv, andload.pyputs both into the database. - The website displays the publication with an A* badge and the metrics in the table columns.
8. Universities Covered
The database covers Accounting & Finance researchers at the eight Group of Eight Australian universities. Each card shows the name stored in the database and the university's output folder.
Scraping is limited to researchers listed under Accounting & Finance on each university's business school research portal.