Performance Evaluation of TabPFN for Student Depression Prediction Across Varying Sample Sizes
DOI:
https://doi.org/10.21609/jiki.v19i2.1716Abstract
Early prediction of depression in students is a critical challenge, often hindered by the scarcity of large, labelled datasets. While supervised tabular classifiers such as Random Forest, XGBoost, and CatBoost are powerful, they typically require sufficient data and careful hyperparameter optimisation (HPO) to generalise effectively. This paper evaluates the Tabular Prior-data Fitted Network (TabPFN), a foundation model for supervised tabular learning, as a zero-shot classifier for student depression prediction. We conduct a comparative study against five robustly configured baseline classifiers (Random Forest, XGBoost, CatBoost, SVM, and Naive Bayes) across three publicly available student mental health datasets of varying sizes sourced from Kaggle: a micro-sample dataset with 101 instances, a small-sample dataset with 7,022 instances, and a moderate-sample dataset with 27,901 instances. Dataset categorization by size is defined relative to TabPFN v2.5's operational capacity of 50,000 samples rather than general machine learning conventions. Using F1-Score as the primary evaluation metric, our empirical results demonstrate a performance crossover linked to data size. On the microsample, imbalanced dataset, TabPFN achieved the highest F1-Score of 0.727, outperforming the best baseline (CatBoost and Random Forest, F1 = 0.667). In the ablation study, both raw and preprocessed inputs yielded identical results for TabPFN on this dataset, highlighting its capacity to handle unprocessed data without performance loss. On the small and moderate datasets, the tuned baselines were competitive or superior, with CatBoost leading on the moderate-sample dataset (F1 = 0.869). We conclude that TabPFN is an effective and efficient baseline for depression prediction tasks in datascarce environments, providing competitive results without HPO, while traditional ensembles remain preferred for larger datasets.
Downloads
Published
How to Cite
Issue
Section
License
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).










