The Agnew Relationship Measure – 5 (ARM-5; Cahill et al., 2012) is a five-item, client-rated measure of the therapeutic alliance, developed as a short form of the 28-item ARM (Agnew-Davies et al., 1998). Completed in under a minute by clients aged 18 and over, it is designed for repeated use in Feedback Informed Treatment, where clinicians gather brief feedback on the working relationship and use it to adjust care while therapy is underway (Miller et al., 2016).
The Agnew Relationship Measure – 5 (ARM-5) is a five-item self-report measure of therapeutic alliance (Cahill et al., 2012). Widely used by clinicians engaging in Feedback Informed Treatment (FIT), it is completed by clients aged 18 and over and designed for repeated use in routine practice, close to the end of a session. Therapeutic alliance, the collaborative bond between client and therapist, is one of the most consistent predictors of psychotherapy outcomes, with stronger alliance linked to better outcomes across therapy types, presenting problems, measures, and countries (Flückiger et al., 2018).
The therapeutic alliance helps therapy outcomes because the driver of the bond is essentially trust (Bordin, 1979), and trust gives the client the safety to open up (Farber, 2003). When clients feel safe with their therapist they disclose more of what matters, take the emotional risks involved in facing difficult material, and stay engaged rather than withdrawing. This appears to be more than a by-product of therapy going well. The therapeutic alliance mediates treatment outcomes across a large body of studies (Baier, Kline & Feeny, 2020), and session-to-session gains in the alliance predict later symptom improvement (Falkenström et al., 2013). A trusting alliance encourages honest disclosure and active engagement through which therapy works, which is why monitoring it, and responding when it weakens, matters.
Because the therapeutic alliance predicts outcomes and can shift over the course of therapy, it is one of the two concepts tracked in FIT. FIT is a structured approach in which clinicians routinely gather brief client feedback on two things, (i) how the client is progressing with regard to symptoms and (ii) the quality of the therapeutic relationship. The clinician would use the feedback to adjust care while therapy is still underway (Miller et al., 2016). Rather than relying on clinical impression alone, the clinician tracks each client’s response from session to session and responds to it directly. Adding routine feedback to treatment has been associated with improved outcomes, with the clearest benefit for clients who are not progressing as expected (Lambert et al., 2018; de Jong et al., 2021).
Monitoring the therapeutic alliance helps detect ruptures: strains or breakdowns in the working relationship that show up either as disagreement over the goals or tasks of therapy, or as a weakening of the bond. While ruptures are common, clients often do not raise them unless asked, so a low or falling ARM-5 score can indicate a rupture that the therapist might otherwise miss. By highlighting a potential rupture, the clinician is able to begin addressing what is causing the rupture and improve client outcomes (Eubanks et al., 2018; Miller et al., 2016). A dip in ARM-5 score is best treated as a prompt to ask what is happening in the relationship rather than as a verdict on it.
The ARM-5 is designed primarily as a single index of overall therapeutic alliance by reducing the scope of the original 28-item scale to three areas:
Monitoring the therapeutic alliance has direct clinical value. A weaker alliance is one of the more reliable warning signs that a client may disengage from treatment (Sharf et al., 2010), so detecting a decline early creates an opportunity to address the relationship before the client leaves therapy. The ARM-5 is suited to regular administration and discussion at the end of a session. Where session-by-session use is impractical, a lower-frequency schedule works just as well; Di Malta and colleagues (2025), for example, administered it after the first session, after the third, and then every fourth session.
Clients tend to rate the alliance highly, so ARM-5 scores cluster near the top of the scale, a ceiling effect common to all alliance measures (Meier, 2022; Cahill et al., 2012). This ceiling effect means that comparing the score to a normative sample is not that useful, as changes in the higher percentiles reflect very small change in the ARM-5. The average score and descriptor bands are the focus of the ARM, particularly scores low enough to suggest disagreement or doubt. In this case an average of less than 6 indicates responses at the “slightly agree” or less, suggesting that the alliance is Challenged or Significantly Challenged. Another key output of the ARM-5 is change over time; a decline, even a small one from a high score, is the main prompt to attend to the relationship (Miller et al., 2016).
The ARM-5 produces a total core alliance score. The total is reported both as a raw score (ranging from 5 to 35) and as an average item score (ranging from 1 to 7), where higher scores indicate a stronger client-rated alliance. Because the three content areas (Bond, Partnership, and Confidence) are highly interrelated and tend to move in the same direction, the total score is the primary output.
The ARM-5’s five items map onto three content areas of the alliance:
The three areas are shown because they add clinical detail the total alone cannot: seeing which aspect of the relationship a client rated lower, whether the bond, the sense of partnership, or their confidence in the therapist and the approach, gives the clinician a concrete place to begin a conversation about repairing or strengthening the alliance.
A percentile is presented for the total alliance score to give relative context. The percentile compares the client’s average alliance score against the NovoPsych routine-care sample. A percentile of 50 marks the median score, the middle of the sample. A higher percentile indicates a stronger client-rated alliance relative to the sample, and a lower percentile indicates a weaker alliance. Because client alliance ratings cluster high, the percentile is best read as a relative position within a therapy sample rather than as an absolute judgement of the relationship, and the highest possible ARM-5 score sits at around the 82nd percentile.
Total ARM-5 scores are grouped into four descriptor bands. The bands are anchored to the response scale, where 6 is “agree” and 7 is “strongly agree”, so an average of 6.6 or above means the client agrees the alliance is strong; the percentile is shown alongside as relative context against the adult reference sample.
Most clients fall in the Moderate or Strong bands; the lower two flag relationships to attend to and openly discuss. A drop in score over time is more telling than any single rating.
The ARM-5 is designed for repeated administration so that change in the alliance can be tracked over the course of therapy. The direction and size of any change in the total score since the previous administration are reported and plotted over time. Any decrease in a client’s score is treated as potentially meaningful and worth attention. Even small, session-to-session drops in the alliance can matter: single-point decreases on a brief alliance measure have been associated with poorer outcomes at the end of treatment, even when scores otherwise remain high (Miller et al., 2016), and session-to-session change in the alliance predicts later symptom change (Falkenström et al., 2013). A decline is therefore highlighted as a prompt to consider and discuss the working relationship.
On a first administration the report presents the client’s current total alliance score against the four descriptor bands, with the percentile shown as context on the right-hand side. Because alliance scores cluster near the top of the scale, the chart’s average-score axis is truncated to a floor of 4 (the neutral point) whenever the score is above 4, so that differences near the ceiling remain visible; if the score is 4 or below, the axis extends to the full 1 to 7 range. A second chart displays the Bond, Partnership and Confidence areas as separate bars.
On repeat administrations a line chart plots the total alliance score across administrations so the trajectory of the working relationship is visible at a glance, with a secondary line chart of the three content areas showing which aspect of the alliance is shifting. The area bars and trajectories are there to aid discussion. Where the current alliance is low, or has declined since the previous administration, the report flags this as a prompt to attend to the alliance.
The ARM-5 is a five-item short form of the 28-item Agnew Relationship Measure (Agnew-Davies et al., 1998), developed by Cahill and colleagues (2012) using Rasch analysis to capture the core alliance (bond, partnership, and confidence) in as few items as possible, and provides a single score of alliance strength. Its reliability is strong, with internal consistency around .79 to .85 across depression and general psychotherapy samples (Cahill et al., 2012; Di Malta et al., 2025). Validity is well supported: the parent measure correlates very strongly with the most widely used alliance instrument, the Working Alliance Inventory (Stiles et al., 2002), and higher alliance on the ARM and its short forms predicts greater symptom improvement by the end of treatment (Stiles et al., 1998; Cahill et al., 2012; Peji & Perez, 2026).
The three areas measured by the ARM-5 (Bond, Partnership, Confidence) overlap so closely that they essentially measure the same thing, and they could not be reliably told apart when the short form was developed (Cahill et al., 2012). The total core alliance score is therefore the main result, and the three areas are shown for completeness rather than as separate scores to interpret.
Percentiles are derived empirically from a NovoPsych routine-care sample of 2,674 adult clients (aged 18 and over) who completed the ARM-5 on first administration, between January 2019 and October 2025. Data was drawn from 30 clinicians who had opted to have their de-identified data used to improve clinical scales. Clients were 60.5% female, 30.9% male, 0.3% gender diverse or other, and 8.3% not specified, and ranged in age from 18 to 82 years (M = 36.9; SD = 13.4). Most clients (52.3%) completed the ARM-5 on multiple occasions, while 47.7% completed it once; among those assessed more than once, the average number of administrations was 4.26 (SD = 2.99).
In this sample the average response on first administration was 6.42 (SD = 0.69; median = 6.6), equivalent to a total score of 32.09 (SD = 3.45; median = 33). As is typical of alliance measures (Meier, 2022; Cahill et al., 2012), ratings are clustered at the high end of the scale, giving a marked ceiling effect. Percentiles are therefore calculated empirically from the observed distribution.
The four descriptor bands are derived by NovoPsych and anchored to the meaning of the ARM-5 response scale rather than to percentile landmarks. On the 1-to-7 scale a rating of 6 corresponds to “agree” and 7 to “strongly agree”, so that an average of 6 or above reflects a client who, overall, agrees the alliance is strong. An average below 6 requires some responses to be “slightly agree” or weaker: even a score that looks high, such as 5.8, indicates a relationship the client is not yet fully endorsing and is classed as Challenged, while averages below 5.0 indicate a more substantial concern.
This anchoring is a deliberate response to the ceiling effect. Because client ratings cluster near the top of the scale, percentiles compress at the upper end (the maximum average score sits at only around the 82nd percentile), so bands based on percentile alone would draw their main distinctions among near-maximum scores and tell clinicians little. Anchoring to the response-scale meaning sets the bands deliberately high: most clients fall in Moderate or Strong, and the lower two bands flag the minority of relationships that warrant attention. Percentiles are retained as relative context rather than as the basis for the cut-points.
The therapeutic alliance describes the quality of the working relationship between a client and their therapist: the sense of a supportive bond, agreement about the goals and tasks of therapy, and the client’s confidence in the therapist and the approach. Decades of research show that a stronger alliance is one of the most reliable predictors of better outcomes across different types of therapy. Measuring it gives the therapist direct feedback on the relationship itself, which is otherwise easy to misjudge.
A declining score indicates that, from the client’s perspective, the working relationship has weakened since a previous administration. Because a client’s own change over time is more informative than the absolute level, a drop is the signal to attend to. Alliance ratings can shift for many reasons, including a difficult session, a misunderstanding, or a change in how therapy is going, so a drop is best read as a signal about the current state of the relationship rather than a fixed judgement, and as a prompt to discuss and repair. Eliciting honest feedback, including some negative feedback, gives the clinician the chance to address concerns before the client disengages.
Clients tend to rate the alliance positively, so scores cluster near the top of the scale and the average is often quite high. A change over time, particularly a drop, is often more informative than the absolute level at any single point.
Because it is so brief, the ARM-5 is designed to be completed regularly rather than as a one-off; often as frequently as each session, so the working relationship can be checked in on as therapy unfolds. The value comes from a client’s own pattern over time rather than any single rating, so it helps to administer it consistently and to have the client anchor each response to the same therapeutic relationship. Even a single administration can open a useful conversation about how the work feels from the client’s side.
The ARM-5 and the Session Rating Scale (SRS) are both very brief, client-rated measures of the therapeutic alliance, and both are widely used in feedback-informed treatment to check in on the working relationship from session to session. They overlap a great deal: both rest on the same underlying model of the alliance, an emotional bond together with agreement on the goals and tasks of therapy, and the ARM-5 has been shown to predict treatment outcomes to a similar degree as the SRS. For session-to-session alliance monitoring the two are largely interchangeable.
Clinicians are not always well placed to judge how therapy is going from impression alone, and clients often will not volunteer that the work feels off track or that the relationship has strained. FIT closes that gap by building a short, regular feedback loop into treatment: the client rates their progress and the working relationship, and the clinician sees the result straight away and can act on it in the same session. Its main value is catching the clients who are quietly not improving, or beginning to disengage, early enough to change course, whether by adjusting the approach, repairing a strain in the relationship, or raising something the client had not felt able to name. Routinely gathering this feedback has been associated with better outcomes and lower dropout, with the clearest benefit for clients who are not progressing as expected. It also gives the client an explicit voice in their own care, which supports the collaboration that drives good outcomes.
Cahill, J., Stiles, W. B., Barkham, M., Hardy, G. E., Stone, G., Agnew-Davies, R., & Unsworth, G. (2012). Two short forms of the Agnew Relationship Measure: The ARM-5 and ARM-12. Psychotherapy Research, 22(3), 241–255. https://doi.org/10.1080/10503307.2011.643253
Agnew-Davies, R., Stiles, W. B., Hardy, G. E., Barkham, M., & Shapiro, D. A. (1998). Alliance structure assessed by the Agnew Relationship Measure (ARM). British Journal of Clinical Psychology, 37(2), 155–172. https://doi.org/10.1111/j.2044-8260.1998.tb01291.x
Baier, A. L., Kline, A. C., & Feeny, N. C. (2020). Therapeutic alliance as a mediator of change: A systematic review and evaluation of research. Clinical Psychology Review, 82, 101921. https://doi.org/10.1016/j.cpr.2020.101921
Baldwin, S. A., Wampold, B. E., & Imel, Z. E. (2007). Untangling the alliance-outcome correlation: Exploring the relative importance of therapist and patient variability in the alliance. Journal of Consulting and Clinical Psychology, 75(6), 842–852. https://doi.org/10.1037/0022-006X.75.6.842
Bordin, E. S. (1979). The generalizability of the psychoanalytic concept of the working alliance. Psychotherapy: Theory, Research & Practice, 16(3), 252–260. https://doi.org/10.1037/h0085885
Cahill, J., Stiles, W. B., Barkham, M., Hardy, G. E., Stone, G., Agnew-Davies, R., & Unsworth, G. (2012). Two short forms of the Agnew Relationship Measure: The ARM-5 and ARM-12. Psychotherapy Research, 22(3), 241–255. https://doi.org/10.1080/10503307.2011.643253
de Jong, K., Conijn, J. M., Gallagher, R. A. V., Reshetnikova, A. S., Heij, M., & Lutz, M. C. (2021). Using progress feedback to improve outcomes and reduce drop-out, treatment duration, and deterioration: A multilevel meta-analysis. Clinical Psychology Review, 85, 102002. https://doi.org/10.1016/j.cpr.2021.102002
Di Malta, G., Vos, J., Bond, J., van Rijn, B., & Cooper, M. (2025). Relational depth as predictor of mental health outcomes in psychotherapy: Random intercept cross-lagged analyses. Psychotherapy Research. Advance online publication. https://doi.org/10.1080/10503307.2025.2551073
Eubanks, C. F., Muran, J. C., & Safran, J. D. (2018). Alliance rupture repair: A meta-analysis. Psychotherapy, 55(4), 508–519. https://doi.org/10.1037/pst0000185
Falkenström, F., Granström, F., & Holmqvist, R. (2013). Therapeutic alliance predicts symptomatic improvement session by session. Journal of Counseling Psychology, 60(3), 317–328. https://doi.org/10.1037/a0032258
Farber, B. A. (2003). Patient self-disclosure: A review of the research. Journal of Clinical Psychology, 59(5), 589–600. https://doi.org/10.1002/jclp.10161
Flückiger, C., Del Re, A. C., Wampold, B. E., & Horvath, A. O. (2018). The alliance in adult psychotherapy: A meta-analytic synthesis. Psychotherapy, 55(4), 316–340. https://doi.org/10.1037/pst0000172
Lambert, M. J., Whipple, J. L., & Kleinstäuber, M. (2018). Collecting and delivering progress feedback: A meta-analysis of routine outcome monitoring. Psychotherapy, 55(4), 520–537. https://doi.org/10.1037/pst0000167
Meier, S. T. (2022). Investigation of causes of ceiling effects on working alliance measures. Frontiers in Psychology, 13, 949326. https://doi.org/10.3389/fpsyg.2022.949326
Miller, S. D., Bargmann, S., Chow, D., Seidel, J., & Maeschalck, C. (2016). Feedback-informed treatment (FIT): Improving the outcome of psychotherapy one person at a time. In W. O’Donohue & A. Maragakis (Eds.), Quality improvement in behavioral health (pp. 247–262). Springer. https://doi.org/10.1007/978-3-319-26209-3_16
Peji, R. G. A., & Perez, J. A. (2026). Therapeutic alliance, treatment intensity, and symptom change in post-traumatic stress disorder: A retrospective study of eye movement desensitization and reprocessing-centered psychotherapy. Cureus, 18(2), e103603. https://doi.org/10.7759/cureus.103603
Sharf, J., Primavera, L. H., & Diener, M. J. (2010). Dropout and therapeutic alliance: A meta-analysis of adult individual psychotherapy. Psychotherapy: Theory, Research, Practice, Training, 47(4), 637–645. https://doi.org/10.1037/a0021175
Stiles, W. B., Agnew-Davies, R., Hardy, G. E., Barkham, M., & Shapiro, D. A. (1998). Relations of the alliance with psychotherapy outcome: Findings in the Second Sheffield Psychotherapy Project. Journal of Consulting and Clinical Psychology, 66(5), 791–802. https://doi.org/10.1037/0022-006X.66.5.791
Stiles, W. B., Agnew-Davies, R., Barkham, M., Culverwell, A., Goldfried, M. R., Halstead, J., Hardy, G. E., Raue, P. J., Rees, A., & Shapiro, D. A. (2002). Convergent validity of the Agnew Relationship Measure and the Working Alliance Inventory. Psychological Assessment, 14(2), 209–220. https://doi.org/10.1037/1040-3590.14.2.209