How does GraalVM make the optimization in Reflexion possible
My previous article explained how Reflexion implements this optimization. In this article, I’ll explain how GraalVM makes that optimization possible.
To verify that GraalVM actually performs these optimizations, we can inspect the generated machine code.
Luckily, there is a way to read the assembly generated by Graal Compiler, allowing us to inspect the machine code that Graal actually produces.
The key question is whether Graal Compiler is able to inline the reflective invocation all the way down to the target method. If it succeeds, very little of the original abstraction should remain in the generated machine code.
First, create a compiler directives file:
1
2
3
4
5
6
7
[
{
match: "*::*Call*",
PrintAssembly: true,
PrintInlining: true
}
]
Then run the benchmark with the following VM options: -XX:+UnlockDiagnosticVMOptions -XX:PrintAssemblyOptions=intel -XX:CompilerDirectivesFile=directives.json, we can inspect the generated assembly code.
For Reflexion custom call, the remaining work is simply checking that the cached object is the expected lambda class and that noInstance is false after optimization.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
// Load the compressed oop.
0x00000209a30a6710: mov ebx,DWORD PTR [r9+0xac] ; ImmutableOopMap {rdx=Oop rbx=NarrowOop r8=Oop r9=Oop r10=Oop r11=Oop [40]=Oop [48]=Oop [56]=Oop }
;*aload {reexecute=1 rethrow=0 return_oop=0}
; - (reexecute) io.github.lumine1909.jmh_generated.Benchmark_customCall_jmhTest::customCall_avgt_jmhStub@13 (line 190)
// Read a boolean, the noInstance field.
0x00000209a30a6717: mov edi,DWORD PTR [r12+rbx*8+0x24]; implicit exception: dispatches to 0x00000209a30a74e8
0x00000209a30a671c: movsx ebx,BYTE PTR [r12+rbx*8+0x14]; ImmutableOopMap {rdx=Oop r8=Oop r9=Oop r10=Oop r11=Oop [40]=Oop [48]=Oop [56]=Oop }
;*aload {reexecute=1 rethrow=0 return_oop=0}
; - (reexecute) io.github.lumine1909.jmh_generated.Benchmark_customCall_jmhTest::customCall_avgt_jmhStub@13 (line 190)
// Compare the object's Klass pointer, check if that object is a specific Lambda class.
0x00000209a30a6722: cmp DWORD PTR [r12+rdi*8+0x8],0x105f170
; implicit exception: dispatches to 0x00000209a30a750d
; {metadata('io/github/lumine1909/reflexion/internal/MethodHolder$$Lambda+0x000000001905f170')}
0x00000209a30a672b: jne 0x00000209a30a7225 ;*invokeinterface get {reexecute=0 rethrow=0 return_oop=0}
; - io.github.lumine1909.reflexion.Method::invoke@11 (line 141)
; - io.github.lumine1909.Benchmark::customCall@11 (line 169)
; - io.github.lumine1909.jmh_generated.Benchmark_customCall_jmhTest::customCall_avgt_jmhStub@17 (line 190)
// Check if field noInstance is false (the method is a virtual method).
0x00000209a30a6731: test bl,bl
0x00000209a30a6733: jne 0x00000209a30a71fe ;*ifeq {reexecute=0 rethrow=0 return_oop=0}
; - io.github.lumine1909.reflexion.Method::invoke@31 (line 142)
; - io.github.lumine1909.Benchmark::customCall@11 (line 169)
; - io.github.lumine1909.jmh_generated.Benchmark_customCall_jmhTest::customCall_avgt_jmhStub@17 (line 190)
// GC related things.
0x00000209a30a6739: cmp BYTE PTR [r15+0x48],0x0
0x00000209a30a673e: xchg ax,ax
0x00000209a30a6740: je 0x00000209a30a6775
0x00000209a30a6746: mov edi,DWORD PTR [r11+0xc]
0x00000209a30a674a: shl rdi,0x3
0x00000209a30a674e: test rdi,rdi
0x00000209a30a6751: je 0x00000209a30a6775
0x00000209a30a6757: mov rbx,QWORD PTR [r15+0x38]
0x00000209a30a675b: cmp rbx,0x0
0x00000209a30a675f: nop
0x00000209a30a6760: je 0x00000209a30a732f
0x00000209a30a6766: sub rbx,0x8
0x00000209a30a676a: mov QWORD PTR [r15+0x38],rbx
0x00000209a30a676e: add rbx,QWORD PTR [r15+0x40]
0x00000209a30a6772: mov QWORD PTR [rbx],rdi
// What the target method actually does (set a String field to "42").
0x00000209a30a6775: mov DWORD PTR [r11+0xc],0xe2585d69; {oop("42"{0x0000000712c2eb48})}
In contrast, mhCall (MethodHandle invokeExact) is not fully inlined.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
// Attempt further inlining.
0x0000023e70856e6f: call 0x0000023e7024f3e0 ; ImmutableOopMap {[32]=Oop [40]=Oop [48]=Oop [56]=Oop [64]=Oop [72]=Oop [80]=Oop }
;*invokestatic maybeCustomize {reexecute=0 rethrow=0 return_oop=0}
; - java.lang.invoke.Invokers::checkCustomized@19 (line 626)
; - java.lang.invoke.LambdaForm$MH/0x0000000035095400::invokeExact_MT@14
; - io.github.lumine1909.Benchmark::mhCall@8 (line 139)
; - io.github.lumine1909.jmh_generated.Benchmark_mhCall_jmhTest::mhCall_avgt_jmhStub@17 (line 190)
; {static_call}
0x0000023e70856e74: nop DWORD PTR [rax+rax*1+0x2800076c]; {post_call_nop}
0x0000023e70856e7c: mov rdx,QWORD PTR [rsp+0x20]
0x0000023e70856e81: mov r8,QWORD PTR [rsp+0x30]
// MethodHandle invoke the target through invokeBasic, introducing an additional call.
0x0000023e70856e86: mov r9d,0x2a ;*invokevirtual invokeBasic {reexecute=0 rethrow=0 return_oop=0}
; - java.lang.invoke.LambdaForm$MH/0x0000000035095400::invokeExact_MT@20
; - io.github.lumine1909.Benchmark::mhCall@8 (line 139)
; - io.github.lumine1909.jmh_generated.Benchmark_mhCall_jmhTest::mhCall_avgt_jmhStub@17 (line 190)
0x0000023e70856e8c: data16 xchg ax,ax
0x0000023e70856e8f: call 0x0000023e7024ede0 ; ImmutableOopMap {[40]=Oop [48]=Oop [56]=Oop [64]=Oop [72]=Oop [80]=Oop }
;*invokevirtual invokeBasic {reexecute=0 rethrow=0 return_oop=1}
; - java.lang.invoke.LambdaForm$MH/0x0000000035095400::invokeExact_MT@20
; - io.github.lumine1909.Benchmark::mhCall@8 (line 139)
; - io.github.lumine1909.jmh_generated.Benchmark_mhCall_jmhTest::mhCall_avgt_jmhStub@17 (line 190)
; {optimized virtual_call}
The generated machine code shows that Graal Compiler completely inlines Reflexion’s Method invocation, while the MethodHandle version still goes through invokeBasic. This difference is consistent with the benchmark results from the previous article.