Open-access Psychometric properties of the Stroop-App in the remote assessment of children/adolescents

Propriedades psicométricas do Stroop-App na avaliação remota de crianças e adolescentes

Propiedades psicométricas de la Stroop-App en la evaluación remota de niños/adolescentes

Abstract

The study investigated the psychometric properties of the Stroop-App, applied remotely to access inhibitory control. 128 children and adolescents between 10 and 17 years old participated (general sample), 24 with Attention Deficit Hyperactivity Disorder (ADHD Group), evaluated with Stroop-App and CPT-Flex, via videoconference. Guardians completed scales to assess executive functioning (IFERA-I) and indicators of inattention and hyperactivity/impulsivity (SNAP-IV). The Stroop-App showed high internal consistency. Significant but low relationships were observed between performances on the Stroop-App, CPT-Flex and responses to IFERA-I and SNAP-IV. Compared to the matched control group (taken from the general sample), the ADHD group was slower, with greater variability in response time (RT). Intra-group comparison indicated differential demands of the Stroop-App conditions. The Stroop-App demonstrated accuracy and evidence of validity, suggesting greater utility of RT and variability measures. The advantage of computerization of the instrument and the innovation of remote application constitute aspects of technological contribution.

Keywords:
Neuropsychological assessment; Executive function; Cognitive processes

Resumo

O estudo investigou as propriedades psicométricas do Stroop-App, aplicado remotamente, para avaliar controle inibitório. Participaram 128 crianças e adolescentes entre 10 e 17 anos (amostra geral), 24 com Transtorno de Déficit de Atenção e Hiperatividade (Grupo TDAH), avaliados com Stroop-App e CPT-Flex via videoconferência. Responsáveis preencheram escalas para avaliar funcionamento executivo (IFERA-I) e indicadores de desatenção e hiperatividade/impulsividade (SNAP-IV). Stroop-App mostrou alta consistência interna. Relações significativas, mas baixas, foram observadas entre os desempenhos no Stroop-App, CPT-Flex e as respostas ao IFERA-I e SNAP-IV. Comparado ao grupo controle pareado (extraído da amostra geral), o grupo TDAH foi mais lento, com maior variabilidade no tempo de resposta (TR). Comparação intra-grupos indicou demandas diferenciais das condições do Stroop-App. O Stroop-App demonstrou precisão e evidências de validade, sugerindo maior utilidade das medidas de TR e variabilidade. A vantagem da computadorização do instrumento e a inovação da aplicação remota constituem aspectos de contribuição tecnológica.

Palavras-chave:
Avaliação neuropsicológica; Função executiva; Processos cognitivos

Resumen

El estudio investigó las propiedades psicométricas de el Stroop-App, aplicado de forma remota, para evaluar el control inhibitorio. Participaron 128 niños y adolescentes entre 10 y 17 años (muestra general), 24 con Trastorno por Déficit de Atención e Hiperactividad (TDAH), evaluados con Stroop-App y CPT-Flex vía videoconferencia. Los tutores completaron escalas para evaluar funcionamiento ejecutivo (IFERA-I) y inatención e hiperactividad/impulsividad (SNAP-IV). Stroop-App mostró una alta coherencia interna. Se observaron relaciones significativas bajas entre desempeño en Stroop-App, CPT-Flex y las respuestas a IFERA-I y SNAP-IV. En comparación con el grupo de control emparejado, el grupo con TDAH fue más lento y presentó mayor variabilidad en el tiempo de respuesta (TR). La comparación intragrupo indicó demandas diferenciales en las condiciones de Stroop-App. Stroop-App demostró precisión y validez, sugiriendo mayor utilidad de las medidas de TR y variabilidad. La ventaja de la informatización y la innovación de la aplicación remota constituyen contribución tecnológica.

Palabras clave:
Evaluación neuropsicológica; Función ejecutiva; Procesos cognitivos

Inhibitory control (IC) has been frequently and consistently considered one of the core functions that make up the broader construct of executive functions (EF) (Baggetta & Alexander, 2016; Diamond, 2013; Karr et al., 2018; Miyake & Friedman, 2012). While EF are controlled processes involved in the conscious control of behavior, emotions, and cognition (Dias & Malloy-Diniz, 2023; Zelazo, 2020), IC encompasses response inhibition (i.e., self-control or inhibition of behavior) and interference control, which includes both cognitive inhibition (e.g., suppression of memories and automatic thoughts) and attentional inhibition (e.g., eliminating distractions and focusing on relevant information and stimuli) (Diamond, 2013). According to the author, without IC, humans would be “creatures of habit.”

Among the classic paradigms in Neuropsychology used to for assess IC, the Stroop paradigm stands out. Created in the 1930s, it encompasses the idea that information takes longer to process when there is interference that generates cognitive incongruence. Usually, in analogy to the original version, three cards are used: in the first, the names of colors written in black ink must be read; in the second, the colors in colored elements must be named; in the third, the color of the ink in which a word is written must be stated, inhibiting its reading. The cognitive incongruence generated in this step is known as the “Stroop Effect” (throughout this article, versions similar to this, with demands for reading and naming colors, will be referred to as Stroop Color-Word). Several versions have been created, adapted, and used to assess attention and IC (Coutinho et al., 2018).

Studies using the Stroop Color-Word task versions with samples of children and adolescents suggest an increase in IC, especially during late childhood, with stability in adolescence (Dias, 2009; Leon-Carrion, 2004), highlighting the impact of word reading on performance (without reading automation, there is no cognitive interference) (Leon-Carrion, 2004). Such results are influenced by the version and types of instrument used. For example, no performance progression was observed in a sample of 9 to 15 years old, using Stroop Victoria (a Color-Word version) and Stroop Happy-Sad (a wordless version that uses drawings as stimuli) (Zanini et al., 2021), unlike the studies by Leon-Carrion and Dias, which may have resented more sensitive indices for having used computerized versions. Zanini et al. (2021), on the other hand, identified an effect of the test condition, i.e., participants tend to make more mistakes and be slower, demonstrating “executive cost,” in the incongruent parts of the task versions.

Studies have also pointed out that the relations between indices in the Stroop tasks, in their different versions, with other EF measures (as well as between other EF measures among themselves) tend to be weak, at most moderate (Arán Filippetti et al., 2021; Dias, 2009; Lin et al., 2019; Zanini et al., 2021). The relation between the same index (inhibition cost - interference between task conditions) of two Stroop task versions (Victoria and Happy-Sad) was only 0.24 (Zanini et al., 2021). This may be due to other demands present in the tasks, and then the shared variance is limited, which should be considered in studies on these measures.

The widespread use of instruments based on the Stroop paradigm is confirmed in reviews of instruments used in the field (Ramos & Hamdan, 2016; Santana et al., 2019), including in adolescents, where tasks based on the paradigm were among the five most used tasks (Nyongesa et al., 2019). In spite of the widespread use of that measure and other ones, reviews highlight the insufficiency of psychometric studies in the area. In the national literature, a specific review of the paradigm identified different versions, presentation formats, and types of stimuli, with a wide diversity of indices and scores derived from the measures. It highlighted the limitation of psychometric investigations (available in approximately 10% of the studies) and the absence of norms (2% of the studies had norms) (Martins et al., 2023). Another review - this one of impulsivity and IC in adolescents -, identified the use of the Stroop task in six of the 13 studies analyzed. The authors denounce the relative scarcity of studies on these constructs in this age group and reinforce the usefulness of tasks based on the Stroop paradigm in their assessment throughout adolescence (Willhelm et al., 2016).

In this sense, given that instruments based on this classic paradigm are useful and informative, an effort to standardize them and investigate their psychometric qualities is urgently needed. In addition to this effort to develop, adapt, and investigate these measures, one recognizes that the computerization can bring advantages to the neuropsychological testing process, enabling, for example, the collection of more precise information, such as reaction time (RT), in addition to other aspects of standardization and automation. Another advantage is the possibility of remote application. Despite the evidence of the feasibility and good acceptance of remote assessment, the challenge is the limited availability of instruments for this context (Miller & Barr, 2017). In Brazil, there are some initiatives in this direction, such as the EF Solutions platform (Malloy-Diniz et al., 2020) and the SAFE (Soluções para Avaliação das Funções Executivas) platform, from the Vetor publishing house.

In the investigation of psychometric properties of instruments, it is possible to make comparisons between groups to test hypotheses about the validity of an instrument. In the case of an EF instrument, more specifically IC, one possible group for comparison criteria is Attention Deficit Hyperactivity Disorder (ADHD). Executive deficit is not exclusive to ADHD (Zelazo, 2020), nor is it the only alteration present in the condition (Wagner et al., 2016), nor is it a condition for its diagnosis (Faraone et al., 2021). However, there is a frequent association between executive difficulties, including IC, and ADHD (Pievsky & McGrath, 2018; Ramos-Galarza & Pérez-Salas, 2017). In view of this, in this study, the diagnosis of ADHD will be the criterion applied to analyze the Stroop task version used.

Thus, the study aims to analyze the psychometric properties of a computerized Stroop task version, the Stroop-App, in the assessment of children and adolescents, applied remotely. More specifically, the objectives include analyzing the instrument accuracy based on its internal consistency, and validity evidence based on its relation with other variables, including relations with instruments measuring related constructs, relations with executive functioning measure by observer-report and with indicators of inattention and hyperactivity, as well as from comparison with a clinical group with ADHD (criterion). Finally, validity evidence will be verified based on the response process.

Methodology

Participants

Initially, 181 children and adolescents (76.2% girls) were selected to participate, students from the 5th grade of Elementary School (EF) to the 3rd grade of High School (EM) from public and private schools in all regions of Brazil. Of these, 36 had a diagnosis of ADHD made by a doctor, neuropsychologist, or multidisciplinary team, proven by a report. The following exclusion criteria were applied: in the general sample, the presence of any neuropsychiatric condition (except for ADHD; 38 occurrences) and a score above 23 (<95% correct answers) on the Stroop-App part 1, which could suggest some reading difficulty (six occurrences); and in the ADHD group, the presence of known comorbidities (10 occurrences), resulting in 53 exclusions.

The final sample (referred to as the general sample) consisted of 128 participants, 71.9% female and 55.5% from public schools, students from the EF 5th grade to the EM 3rd grade, aged between 10 and 17 years. Of these, 70% reported a family income between R$1,100.00 and R$5,500.00. All Brazilian regions continued to be represented, despite the predominance of the Southeast region (27% of participants were from the state of São Paulo alone). From the general sample, 24 participants had a diagnosis of ADHD (ADHD group) and 104 comprised the non-clinical sample (Table 1).

Table 1
Study participant characterization

The non-clinical sample and the ADHD group were equivalent in terms of distribution by school type (x2=3.931; p=0.415), history of grade repetition (x2=0.469; p=0.615), difficulties in motor and linguistic development (x2=5.693; p=0.058), and family income (x2=3.125; p=0.537). There was a difference in the distribution of participants based on school grade (x2=14.958; p=0.037) and age (t=-3.253; p=0.001), with younger ADHD participants concentrated in the initial school levels; and sex (x2=4.582; p=0.032), with the ADHD group having proportionally more boys. Thus, specifically for the between-group comparison analysis, a non-clinical group matched to the ADHD group was extracted from the non-clinical sample. The matching considered the variables, in this order of importance: age, school level, sex, and type of school. Table 1 presents both groups, non-clinical and ADHD. As would be expected from the matching, the groups are equivalent in terms of distribution in school grades (x2=1.929, p=0.926), type of school (x2=2.980, p=0.395), sex (x2=0.083, p=0.773), and age (t=-0.256; p=0.799). The ADHD group presented a greater indication of symptoms of inattention (U=84.00; p<0.001) and hyperactivity/impulsivity compared to the matched non-clinical group (U=137.00; p=0.002).

Instruments

  • Parent Identification Questionnaire (Qp): Used to gather information about health history, education, socioeconomic level, and geographic location, allowing for sample characterization and exclusion criteria identification.

  • Difficulties in Executive Functions, Regulation and Delay Aversion Inventory - Version for children and adolescents (IFERA-I;Trevisan et al., 2022): It measures behavioral manifestations of difficulties in EF, based on reports from caregivers and/or teachers. It was developed based on the neuropsychology of ADHD and encompasses the constructs of EF, delay aversion, and state regulation. Details regarding its development and the theoretical models considered can be found in Trevisan et al. (2022). It contains 28 items, in 5 sub-scales: working memory (WM; six items), inhibitory control (IC; six items), flexibility (FLEX; five items), delay aversion (DA; five items), and state regulation (SR; six items). Each item is composed of a 1-5 Likert scale, indicating the frequency of difficulties. Higher scores denote greater difficulties. Psychometric data are presented in Trevisan et al. (2022).

  • SNAP-IV (Mattos et al., 2016): Public domain scale, developed based on DSM-IV criterion A for ADHD, adapted for Brazil by Mattos et al. (2016), it consists of 18 items related to symptoms of inattention and hyperactivity/impulsivity, for assessment in children and adolescents based on caregivers’ and teachers’ responses. Participants assign a score of 1 to 4 for each item. Higher scores indicate greater difficulty.

  • CPT-Flex (EF-Toolkit;Malloy-Diniz et al., 2020): EF-Toolkit is a platform that includes classic neuropsychological tasks and allows for computerized assessments. One used the CPT-Flex, a test on this platform, which aims to assess IC and cognitive flexibility. In this version, stimuli (geometric shapes) are presented in gray and red, individually, on the screen, for milliseconds. When in gray, participants must respond to all shapes except one. When in red, the rule is reversed. Total score, mean reaction time (RT), and standard deviation (SD) of RT (variability) were used throughout the test.

  • Stroop-App (Seabra et al., 2020): It is a computerized and online version of the Stroop task, used to assess IC, specifically interference control. The participant follows three steps: The first involves reading words to verify reading ability; the second, naming colors (in colored circles); and the third, resolving cognitive conflict by having to name the color in written words, inhibiting automatic reading behavior. The verbalized responses are recorded, and the software performs automatic correction. Performance is measured by correct answers and RT in parts 2 and 3 (part 1 serves only for reading control), in addition to score and interference time (score or RT part 3 minus score or RT part 2). A measure of variability, SD of RT in parts 2 and 3, was also used.

Ethical and data collection procedures

The research was not pre-registered. The project was approved by the Ethics Committee of the Federal University of Santa Catarina, under number CAEE: 43349121.9.0000.0121. This study is a segment of this research. Recruitment was done through research dissemination on the institution’s and partner institutions’ official websites and their respective social networks.

Initially, the participants’ guardians were presented with the Informed Consent Form. They were then directed to answer the Qp, IFERA-I, and SNAP-IV via Google Forms, and the response time was approximately 30-40 minutes. Subsequently, a session was scheduled for CPT-Flex and Stroop-App being applied online and remotely to the participating children and adolescents. In this contact, the clinical group was asked to send the report with the diagnosis of ADHD. In the scheduled evaluation session, before the start of the application, the Informed Assent Form was presented to the volunteers. Instruments were presented to each participant in an individual video call, which lasted approximately 60 minutes. At the beginning of the sessions, one stressed the importance of the participant not being interrupted and being in a comfortable environment with minimal distractions and maximum privacy. Participants received a link to access the computerized tasks, and the administrator monitored via videoconference. The order of the tasks was counterbalanced. Data from the study may be made available, upon request to the last authors, after studies on the instrument quality and its publication.

Data Analysis

Internal consistency was calculated from the Kuder-Richardson (KR) coefficients for scores (dichotomous items) and Cronbach’s Alpha for RT. The indices were calculated for the Stroop-App parts 2 and 3 (since part 1 only controls reading level, and one verified ceiling effect). The Kolmogorov-Smirnov test revealed that there was no normal distribution for most measures. Spearman correlation analysis was used to verify the relation between Stroop-App scores and those of CPT-FLEX, IFERA-I, and SNAP-IV (the general sample was considered here) (interpretation: 0.20>rho<0.49 - weak; 0.40>rho<0.60 - moderate; rho > 0.60 - strong). Mann-Whitney analysis was used to compare Stroop-App performance in the ADHD group and in the matched non-clinical group. Effect sizes were interpreted according to Lenhard and Lenhard (2022): r<0.30 - small; r between 0.3 and 0.5 - medium; r >0.50 - large. Finally, Wilcoxon analysis was used for intra-group comparison of performance in the conditions (parts 2 and 3) of the Stroop-App. A general sample was used, and in an attempt to define differential profiles across conditions, non-clinical and ADHD samples were used separately.

Results

Analysis of internal consistency, for the general sample and by group, is shown in Table 2. In all cases, coefficients were greater than 0.70. In both Stroop-App, the RT measures proved to be more accurate than the scores.

Table 2
Internal consistency for performance, in terms of score and RT, in the Stroop-App

Table 3 presents the Spearman correlation matrix. Some significant relations with CPT-Flex were evidenced, but they tended to be weak. Despite this, they tended to occur with Stroop indices that include executive demands (except for two relations with indices from part 2), and occurred in the expected direction.

Table 3
Relations between Stroop-App and IFERA-I and SNAP-IV indicators and CPT-Flex performance (general sample)

Several relations were observed between Stroop-App indices and IFERA-I indices. Some, although significant, can be considered null. The others relations were weak, but always in the expected direction (negative with score measures and positive with RT measures), that is, lower scores or higher RT (including greater variability) in the Stroop-App were associated with a greater report of difficulties in IFERA-I. In the score measure and in RT, more relations were established with the Stroop-App part 3, which is the part with executive demands. Furthermore, the RT measures showed more associations with the IFERA-I indices than the performances in terms of score. Linked to these two observations, RT in part 3, the interference RT, and the variability measure in part 3 were associated with all IFERA-I indices. The relations with SNAP-IV indicators, in turn, were established only with the Stroop-App RT measures and, more consistently, with the indices that have executive demand (part 3 and interference). Again, all these relations, although significant, were weak.

The comparison between groups, considering the ADHD and the matched non-clinical groups, revealed significant differences in RT measures in both parts of the test. In both, the ADHD group had higher RT. No difference was found in the score measures. In the SD measure, there was a group effect in the Stroop-App part 3, with greater variability presented in the ADHD group. All significant comparisons presented medium effect size (>0.30). Table 4 presents these results.

Table 4
Comparison of Stroop-App performance between the ADHD group and the matched non-clinical group.

The comparison of intra-group performance (Table 5) across conditions (parts 2 and 3) revealed significant differences, with lower performance in part 3 (with executive demand) compared to part 2 (color naming only), both in the general sample and in the non-clinical sample. The same trend occurred in the ADHD group, but did not reach statistical significance. Considering RT, there was a condition effect, with higher RT in part 3, in all samples (general, non-clinical and ADHD samples).

Table 5
Descriptive and inferential intra-group statistics comparing performance across the Stroop-App stages

Discussion

The study aimed to investigate the psychometric properties of the Stroop-App, applied remotely, in the IC assessment in children and adolescents. The internal consistency coefficients found ranged from 0.70 to 0.97, considered adequate. The precision indices tended to be slightly higher in the Stroop-App part 3 compared to part 2, except in the ADHD group, possibly linked to the fact that this group tends to have greater variability in performance (greater oscillation in terms of errors and RT) in part 3 of the test. Furthermore, in all groups, RT measures proved to be more precise (KR and alpha > 0.93) than score measures. Thus, based on the covariance between items, the Stroop-App was able to produce internally consistent performances (Peixoto & Ferreira-Rodrigues, 2019), proving to be a precise measure. The greater accuracy of RT measures is one of the advantages made possible by the computerization of instruments (Miller & Barr, 2017), and it was found that this accuracy was maintained even in remote assessment.

Regarding the validity evidence based on the relation with other variables, some significant relations were found between the Stroop-App and CPT-FLEX indices, all weak (0.20<observed rho<0.40), although in the expected direction. The relations tended to concentrate on the Stroop indices that include executive demand. These results show a tendency for participants with better performance on one test to also perform better on the other, in terms of correct answers, RT, and variability. This result is congruent, since both tests, despite other cognitive demands involved, measure IC aspects. Weak to moderate relations (between 0.20 and 0.50) are sufficient to conclude that there is evidence of validity based on the correlation with other variables when tests that assess related constructs are used (Nunes & Primi, 2010). The most consistent relationships were observed between the variability indices, suggesting that this may be a useful index to use in the analysis of Stroop tasks. This index will be discussed later.

Consistently, studies have found weak to, at most, moderate relations between FE measures, including the Stroop task versions (Arán Filippetti et al., 2021; Dias, 2009; Lin et al., 2019; Zanini et al., 2021), and also between the same indices in different tasks of the Stroop paradigm (Victoria x Happy-Sad; Zanini et al., 2021). The same occurs with tests that assess the same construct (Arán Filippetti et al., 2021), a study that also revealed that the interference score had the lowest number of relations with other EF tests and, when those did occur, they were the least consistent. Such weak and moderate relationships are expected, given the multidimensional EF nature, even in a child and adolescent sample (Arán Filippetti et al., 2021).

Furthermore, while the Stroop-App requires reading, color naming, and interference control, CPT-FLEX involves perceptual-motor speed, sustained attention, and, in terms of executive demands, response inhibition, plus cognitive flexibility in this specific version. Thus, the demands present in each task may explain the relatively limited shared variance between these tasks. In fact, IC is usually understood as including different aspects of processing, such as response inhibition (more required in CPT-FLEX) and interference control (more required in the Stroop paradigm) (Diamond, 2013). This dissociation is corroborated by different patterns of brain activation related to performance in different IC tasks (dos Santos et al., 2023) and explains the relations found.

The relation between the Stroop-App and other EF measures was also verified using an inventory (IFERA-I), which offers a more functional EF measure as reported by guardians. The pattern of relations observed was the same as described previously. One found weak relations, all indicating that worse performance (lower score, higher RT) and greater variability in the Stroop-App are associated with greater reports of executive difficulties. The consistency of these relations for RT in part 3, interference RT, and variability in part 3 stands out, with relationships with all IFERA-I indices. The weak relations corroborate what has already been found in the literature, i.e., that performance tests and report measures are not assessing the same construct, or at the same processing level. Performance tests would be accessing EF more directly, i.e., at the cognitive level, while report/scales measures could access the behavioral manifestation of these abilities or their dysfunction (Godoy, 2024; Soto et al., 2020). The results of this study seem to be aligned with these observations and may suggest that better interference control (cognitive level of analysis) has shared variance with the outcome or more “executive” functioning in daily life (behavioral level of analysis).

Relations of the Stroop-App indices were not established only with the IC sub-scale, but with all IFERA-I domains. This makes sense when considering that IC may play a primary role in executive behavior. For example, avoiding overbearing responses or conflicting information is fundamental to maintaining and implementing goals, which seems to be a demand present in all tasks involving EF (Miyake & Friedman, 2012). In this sense, it is as if interference control is, to some degree, present in any behavior that demands EF (from, for example, the maintenance and implementation of goals).

It is noteworthy that, despite the relations with all scale dimensions, RT in part 3, the interference RT, and the variability in the Stroop-App part 3 had more consistent, although weak, relations with state regulation. Originating from the cognitive-energetic model (Sergeant, 2000), state regulation (SR) refers to a capacity for allocating “cognitive energy” and regulating this activation over time, relating to response efficiency. In the model, this level of state activation plays a role in the EF activation and engagement. Thus, the relations between the Stroop-App and the SR dimension suggest that individuals with greater difficulty in allocating/regulating cognitive resources are less efficient in dealing with cognitive conflict (they make more mistakes and require more time) and exhibit greater response variability (sometimes faster, sometimes slower) in both parts of the test, but especially in the presence of executive demand.

The relations between the Stroop-App and indicators of inattention and hyperactivity were weak, repeating the pattern of previous relations. The relation between RT in the Stroop-App part 2 and reported inattention suggests that individuals with greater complaints of inattention tend to present greater overall response slowness. The other relations suggest that those with greater difficulty in resolving cognitive conflict tend to be assessed as having more inattention and hyperactivity/impulsivity. In an ADHD sample, moderate to strong relations were found between symptoms and difficulties in EF (Silverstein et al., 2018); however, it should be noted that scales were used to assess EF in that study, and the magnitude of the relations may reflect the assessment method. In this study, given that these relations were explored in the general sample, weak relations were hypothesized and show a tendency toward lower efficiency in dealing with cognitive conflict among individuals with more symptoms reported. Even though this analysis did not involve a clinical sample, this result is supported by studies that have corroborated the association between IC and ADHD (Pievsky & McGrath, 2018). The relations also indicated that the greater the presence of indicators, the more irregular the response is. The most consistent relation was between variability in part 3 and inattention, corroborating findings about the usefulness of this type of measure in the ADHD study (Pievsky & McGrath, 2018).

The comparison between the ADHD group and the matched non-clinical group revealed differences in RT in parts 2 and 3, as well as in the variability in the Stroop-App part 3. Differences between groups were expected since the literature generally points to deficits in IC associated with ADHD (Pievsky & McGrath, 2018; Ramos-Galarza & Pérez-Salas, 2017). It should be noted that only the RT measures revealed differences between groups, showing that performance in terms of score is not necessarily impaired in ADHD people and that RT may be, in fact, a more useful measure. It is worth highlighting that differences occurred not only in RT in part 3, but also in part 2 (without evident executive demand), and that there was no difference between groups in the interference RT. This finding suggests that processing speed may also be compromised in such participants. A review by Pievsky and McGrath (2018) revealed altered processing speed in ADHD samples, but the effect size was small. However, the effect exists and it is necessary to consider how speed is assessed, since computerization may allow for more precise measures, such as RT.

A study using a Stroop Color-Word task version with children and adolescents also found no differences between the ADHD group and the control in the interference measure. The Inattentive group performed worse in the reading stage, potentially producing false negatives in the color-word stage. The Combined group performed worse in the color naming and color-word stages. As a result, the interference index appeared biased and was unable to differentiate between groups. The authors concluded that participants with the disorder perform worse in all parts of the task and suggest that performance across the instrument conditions should be taken into account, not just interference indices, as this may reveal aspects of the difficulty profile (Áran Filippetti et al., 2021).

Thus, while there is lower processing efficiency, with a higher RT in part 2, impacting the interference RT, there is also lower efficiency in solving part 3, possibly impacted by a more general processing speed capacity, beyond the IC demanded by the task. Changes in IC in ADHD have been documented in a review, with a moderate effect size (Pievsky & McGrath, 2018), suggesting that IC is frequently compromised in ADHD, although the effect size does not corroborate that it is the central disorder deficit.

This review revealed that the neurocognitive domain with the largest effect size in the comparison between the ADHD group and the control was the RT variability, a finding that corroborated the idea of a difficulty, present in ADHD, in changing the state of brain activation (Pievsky & McGrath, 2018). This index refers to intra-individual variability in response speed or efficiency, i.e., a tendency for fluctuation in ADHD people’s performance. In this study, only the variability in the Stroop-App parte 3 differentiated the groups. This suggests that it is in the face of executive demands that ADHD participants showed greater fluctuation in their performance. This response pattern seems consistent with the hypothesis that the RT variability is an indicator of the efficiency of attentional control processes (Tamm et al., 2012).

Finally, evidence of validity based on the response process was verified. In general, the same pattern was observed in all samples (there was no statistical significance for scores specifically in the ADHD group, possibly due to the group size, but the effect size was larger than in the other groups). Thus, the difference between performance and RT in the Stroop-App conditions 2 and 3 indicated a tendency towards making more errors and greater RT in the condition with executive demand, as expected. Considering the RT measures, which also proved to be more sensitive in other analyses of this study, the effect sizes were all large. This result indicates that the cognitive demand imposed by the task part 3 in fact does differ from part 2, corroborating the notion of different processes involved in the response (hypothetically, the demand on interference control) (Áran Filippetti et al., 2021; Zanini et al., 2021). In this study, the result also suggests that the ADHD group does not suffer greater interference from condition 3 compared to the non-clinical group or the general sample.

In summary, the study showed that the Stroop-App, especially considering the RT indicators, is a measure with satisfactory internal consistency. The study provided evidence of validity for the instrument based on its relation with other variables, and association patterns were identified as expected based on the literature in the field and theoretically grounded, with the test measuring a construct related to indicators of executive functioning at the behavioral level and indicators of inattention and hyperactivity. Regarding the evidence of validity based on the relation with other variables, considering ADHD an external criterion, the instrument was found to be able to discriminate between a group with the diagnosis and a matched control group. Finally, evidence of validity based on the response process was obtained from the differentiation of performances across the task conditions, suggesting distinct cognitive demands involved.

Among the study limitations, it is possible to mention the lack of more control over reading, for example, based on some speed measure. However, this variable was controlled by excluding participants with comorbidities (including dyslexia) or with a success rate<23 (<95%) in the first part of the test. Another point is the small number of ADHD participants, which made it impossible to consider presentation types in the analyses. It is also mentioned that there is a lack of more appropriate measures that would enable the search for convergence patterns, such as other tasks based on the Stroop paradigm.

Despite the limitations, the study points to the Stroop-App as an instrument with potential for assessing IC in children and adolescents, highlighting the need to consider more sensitive and precise measures, potentially more useful, such as RT, an advantage possible with the task computerization. In addition, based on the results, it is suggested that the variability measure should be incorporated into the analyses of the Stroop paradigm. It should be noted that this study reveals good psychometric qualities of this measure in a context still lacking studies and tools, which is remote assessment (Miller & Barr, 2017). In this sense, its potential technological contribution is added to its scientific contribution (regarding the measure qualities and the usefulness of its indicators), providing the area with a computerized measure, with satisfactory applicability and psychometric quality, in remote assessment.

DATA AVAILABILITY STATEMENT

The research data are not available.

Acknowledgments:

CNPq - National Council for Scientific and Technological Development - (NMD and AGS Research Productivity Grant and EKRML Scientific Initiation Grant - PIBIC -, at the time of data collection) and CAPES (MEOM Master’s scholarship, at the time of data collection).

References

  • Arán Filippetti, V., Richaud, M. C., Krumm, G., & Raimondi, W. (2021). Cognitive and socioeconomic predictors of Stroop performance in children and developmental patterns according to socioeconomic status and ADHD subtype. Psychology & Neuroscience, 14(2), 183-206. https://doi.org/10.1037/pne0000224
    » https://doi.org/10.1037/pne0000224
  • Baggetta, P., & Alexander, P. A. (2016). Conceptualization and operationalization of executive function. Mind, Brain, and Education, 10(1), 10-33. https://doi.org/10.1111/mbe.12100
    » https://doi.org/10.1111/mbe.12100
  • Coutinho, G., Mattos, P., Abreu, N. (2018). Atenção. In. L.F. Malloy-Diniz, D. Fuentes, P. Mattos, & N. Abreu (Orgs.), Avaliação neuropsicológica (pp. 83-89). Artmed.
  • Diamond, A. (2013). Executive functions. Annual Review of Psychology, 64, 135-168. https://doi.org/10.1146/annurev-psych-113011-143750
    » https://doi.org/10.1146/annurev-psych-113011-143750
  • Dias, N. M. (2009) Avaliação neuropsicológica das funções executivas: Tendências desenvolvimentais e evidências de validade de instrumentos [Dissertação de Mestrado, Universidade Presbiteriana Mackenzie]. http://dspace.mackenzie.br/handle/10899/22661
    » http://dspace.mackenzie.br/handle/10899/22661
  • Dias, N. M., & Malloy-Diniz, L. F. (2023). Funções executivas: Ponto de partida para compreensão do construto. In N. M. Dias, & L. F. Malloy-Diniz (Orgs.), Tratado das funções executivas: Modelos teóricos, desenvolvimento e construtos associados (pp. 17-31). Ampla.
  • Godoy, S. (2024). Medidas indiretas (baseadas em escalas) de funções executivas. In N. M. Dias, & L. F. Malloy-Diniz (Orgs.), Tratado das funções executivas: Avaliação e intervenção (Vol. 2, pp. 37-49). Ampla.
  • Karr, J. E., Areshenkoff, C. N., Rast, P., Hofer, S. M., Iverson, G. L., & Garcia-Barrera, M. A. (2018). The unity and diversity of executive functions: A systematic review and re-analysis of latent variable studies. Psychological Bulletin, 144(11), 1147-1185. https://doi.org/10.1037/bul0000160
    » https://doi.org/10.1037/bul0000160
  • Lenhard, W., & Lenhard, A. (2022). Computation of effect sizes. Psychometrica. https://www.psychometrica.de/effect_size.html http://dx.doi.org/10.13140/RG.2.2.17823.92329
    » https://doi.org/10.13140/RG.2.2.17823.92329» https://www.psychometrica.de/effect_size.html
  • Leon-Carrion, J., García-Orza, J., & Pérez-Santamaría, F. J. (2004). Development of the inhibitory component of the executive functions in children and adolescents. The International Journal of Neuroscience, 114(10), 1291-1311. https://doi.org/10.1080/00207450490476066
    » https://doi.org/10.1080/00207450490476066
  • Lin, B., Liew, J., & Perez, M. (2019). Measurement of self-regulation in early childhood: Relations between laboratory and performance-based measures of effortful control and executive functioning. Early Childhood Research Quarterly, 47, 1-8. https://doi.org/10.1016/j.ecresq.2018.10.004
    » https://doi.org/10.1016/j.ecresq.2018.10.004
  • Malloy-Diniz, L. F., Timóteo, A., Serpa, A., & Querino, E. (2020). EF Solutions. https://metacognitiv.com/ef-solutions
    » https://metacognitiv.com/ef-solutions
  • Martins, M. E. O., Tosi, C. M. G., Luz, B. P., Toresan, L. H., de Carvalho, C. F., & Dias, N. M. (2023). O Paradigma de Stroop nos estudos brasileiros: uma revisão de escopo. Psicologia: Teoria e Prática, 25(2), ePTPCP14766-ePTPCP14766. https://doi.org/10.5935/1980-6906/ePTPCP14766.en
    » https://doi.org/10.5935/1980-6906/ePTPCP14766.en
  • Mattos, P., Pinheiro, M. A., Rohde, L., & Pinto, D. (2006). Apresentação de uma versão em português para uso no Brasil do instrumento MTA-SNAP-IV de avaliação de sintomas de transtorno do déficit de atenção/hiperatividade e sintomas de transtorno desafiador e de oposição. Revista de Psiquiatria do Rio Grande do Sul, 28(3), 290-297. https://doi.org/10.1590/S0101-81082006000300008
    » https://doi.org/10.1590/S0101-81082006000300008
  • Miller, J. B., & Barr, W. B. (2017). The technology crisis in neuropsychology. Archives of Clinical Neuropsychology: The Official Journal of the National Academy of Neuropsychologists, 32(5), 541-554. https://doi.org/10.1093/arclin/acx050
    » https://doi.org/10.1093/arclin/acx050
  • Miyake, A., & Friedman, N. P. (2012). The nature and organization of individual differences in executive functions: Four general conclusions. Current Directions in Psychological Science, 21(1), 8-14. https://doi.org/10.1177/0963721411429458
    » https://doi.org/10.1177/0963721411429458
  • Nunes, C. H., & Primi, R. (2010). Aspectos técnicos e conceituais da ficha de avaliação dos testes psicológicos. In Conselho Federal de Psicologia (Org.), Avaliação psicológica: Diretrizes na regulamentação da profissão (pp. 101-128). Conselho Federal de Psicologia.
  • Nyongesa, M. K., Ssewanyana, D., Mutua, A. M., Chongwo, E., Scerif, G., Newton, C. R. J. C., & Abubakar, A. (2019). Assessing executive function in adolescence: A scoping review of existing measures and their psychometric robustness. Frontiers in Psychology, 10. https://doi.org/10.3389/fpsyg.2019.00311
    » https://doi.org/10.3389/fpsyg.2019.00311
  • Peixoto, E. M., & Ferreira-Rodrigues, C. F. (2019). Propriedades psicométricas dos testes psicológicos. In M. N. Baptista, M. Muniz, C. T. Reppold, C. H. S. da S. Nunes, L. de F. Carvalho, R. Primi, ... L. Pasquali (Orgs.), Compêndio de avaliação psicológica (pp. 29-39). Editora Vozes.
  • Pievsky, M. A., & McGrath, R. E. (2018). The neurocognitive profile of attention-deficit/hyperactivity disorder: A review of meta-analyses. Archives of Clinical Neuropsychology, 33(2), 143-157. https://doi.org/10.1093/arclin/acx055
    » https://doi.org/10.1093/arclin/acx055
  • Ramos, A. A., & Hamdan, A. C. (2016). O crescimento da avaliação neuropsicológica no Brasil: Uma revisão sistemática. Psicologia: Ciência e Profissão, 36(2), 471-485. https://doi.org/10.1590/1982-3703001792013
    » https://doi.org/10.1590/1982-3703001792013
  • Ramos-Galarza, C. & Pérez-Salas, C. (2017). Control inhibitorio y monitorización em población infantil con TDAH. Avances en Psicología Latinoamericana, 35(1), 117-130. https://doi.org/10.12804/revistas.urosario.edu.co/apl/a.4195
  • Santana, A. N. de, Melo, M., & Minervino, C. A. (2019). Instrumentos de avaliação das funções executivas: Revisão sistemática dos últimos cinco anos. Revista Avaliação Psicológica, 18(1), 96-107. https://doi.org/10.15689/ap.2019.1801.14668.11
    » https://doi.org/10.15689/ap.2019.1801.14668.11
  • dos Santos, A. J., Machado-Pinheiro, W., Osório, A. A., Seabra, A. G., Teixeira, M. C. T. V., Nascimento, J. de A., & Carreiro, L. R. (2023). Association between ADHD symptoms and inhibition-related brain activity using functional near-infrared spectroscopy (fNIRS). Neuroscience Letters, 792, 136962. https://doi.org/10.1016/j.neulet.2022.136962
    » https://doi.org/10.1016/j.neulet.2022.136962
  • Seabra, A. G., Dias, N. M., & Macedo, E. C. (2020). Stroop App (Software em preparação). Memnon.
  • Sergeant, J. (2000). The cognitive-energetic model: An empirical approach to attention-deficit hyperactivity disorder. Neuroscience and Biobehavioral Reviews, 24(1), 7-12. https://doi.org/10.1016/s0149-7634(99)00060-3
    » https://doi.org/10.1016/s0149-7634(99)00060-3
  • Silverstein, M. J., Faraone, S. V., Leon, T. L., Biederman, J., Spencer, T. J., & Adler, L. A. (2018). The Relationship Between Executive Function Deficits and DSM-5-Defined ADHD Symptoms. Journal of Attention Disorders, 24(1), 108705471880434. https://doi.org/10.1177/1087054718804347
    » https://doi.org/10.1177/1087054718804347
  • Soto, E. F., Kofler, M. J., Singh, L. J., Wells, E. L., Irwin, L. N., Groves, N. B., & Miller, C. E. (2020). Executive functioning rating scales: Ecologically valid or construct invalid? Neuropsychology, 34(6), 605-619. https://doi.org/10.1037/neu0000681
    » https://doi.org/10.1037/neu0000681
  • Tamm, L., Narad, M. E., Antonini, T. N., O’Brien, K. M., Hawk, L. W., Jr., & Epstein, J. N. (2012). Reaction time variability in ADHD: A review. Neurotherapeutics: The Journal of the American Society for Experimental NeuroTherapeutics, 9(3), 500-508. https://doi.org/10.1007/s13311-012-0138-5
    » https://doi.org/10.1007/s13311-012-0138-5
  • Trevisan, B. T., Berberian, A. A., Dias, N. M., Roama-Alves, R. J., & Seabra, A. G. (2022). Development and psychometric properties of the difficulties in executive functions, regulation and delay aversion inventory-version for children and adolescents. Avaliação Psicológica, 21(3), 261-272. http://dx.doi.org/10.15689/ap.2022.2103.19957.02
    » http://dx.doi.org/10.15689/ap.2022.2103.19957.02
  • Wagner, F., Rohde, L. A., & Trentini, C. M. (2016). Neuropsicologia do transtorno de déficit de atenção/hiperatividade: Modelos neuropsicológicos e resultados de estudos empíricos. Psico-USF, 21(3), 573-582. https://doi.org/10.1590/1413-82712016210311
    » https://doi.org/10.1590/1413-82712016210311
  • Willhelm, A. R., Fortes, P. M., Czermainski, F. R., Andrade Rates, A. S., & de Almeida, R. M. M. (2016). Neuropsychological and behavioral assessment of impulsivity in adolescents: A systematic review. Trends in Psychiatry and Psychotherapy, 38(3), 128-135. https://doi.org/10.1590/2237-6089-2015-0019
    » https://doi.org/10.1590/2237-6089-2015-0019
  • Zanini, G. A. V., Miranda, M. C., Cogo-Moreira, H., Nouri, A., Fernández, A. L., & Pompéia, S. (2021). An adaptable, open-access test battery to study the fractionation of executive-functions in diverse populations. Frontiers in Psychology, 12. https://doi.org/10.3389/fpsyg.2021.627219
    » https://doi.org/10.3389/fpsyg.2021.627219
  • Zelazo, P. D. (2020). Executive function and psychopathology: A neurodevelopmental perspective. Annual Review of Clinical Psychology, 16, 431-454. https://doi.org/10.1146/annurev-clinpsy-072319-024242
    » https://doi.org/10.1146/annurev-clinpsy-072319-024242

Edited by

  • Editor:
    Evandro Morais Peixoto.

Publication Dates

  • Publication in this collection
    27 Apr 2026
  • Date of issue
    2026

History

  • Received
    03 June 2024
  • Reviewed
    22 June 2024
  • Accepted
    13 Jan 2025
location_on
Universidade de São Francisco, Programa de Pós-Graduação Stricto Sensu em Psicologia R. Waldemar César da Silveira, 105, Vl. Cura D'Ars (SWIFT), Campinas - São Paulo, CEP 13045-510, Telefone: (19)3779-3771 - Campinas - SP - Brazil
E-mail: revistapsico@usf.edu.br
rss_feed Stay informed of issues for this journal through your RSS reader
Go to top Report error